使用data.table在子组中查找相同的行

时间:2013-05-12 10:52:57

标签: r data.table

我的桌子有两个ID。我希望,对于第一个ID的每个值,找出具有不同第二个ID值的两行是否相同(不包括第二个ID的列......)。 我的表格非常相似(但要小得多):

library(data.table)

DT <- data.table(id   = rep(LETTERS, each=10),
                 var1 = rnorm(260),
                 var2 = rnorm(260))


DT[, id2 := sample(c("A","B"), 10, T), by=id] # I need this to simulate different 
                                              # distribution of the id2 values, for
                                              # each id value, like in my real table

setkey(DT, id, id2)

DT$var1[1] <- DT$var1[2] # this simulates redundances
DT$var2[1] <- DT$var2[2] # inside same id and id2

DT$var1[8] <- DT$var1[2] # this simulates two rows with different id2
DT$var2[8] <- DT$var2[2] # and same var1 and var2. I'm after such rows!

> head(DT, 10)
    id           var1           var2 id2
 1:  A  0.11641260243  0.52202152686   A
 2:  A  0.11641260243  0.52202152686   A
 3:  A -0.46631312530  1.16263285108   A
 4:  A -0.01301484819  0.44273945065   A
 5:  A  1.84623329221 -0.09284888054   B
 6:  A -1.29139503119 -1.90194818212   B
 7:  A  0.96073555968 -0.49326620160   B
 8:  A  0.11641260243  0.52202152686   B
 9:  A  0.86254993530 -0.21280899589   B
10:  A  1.41142798959  1.13666002123   B

我目前正在使用此代码:

res <- DT[, {a=unique(.SD)[,-3,with=F]   # Removes redundances like in row 1 and 2
                                         # and then removes id2 column.
             !identical(a, unique(a))},  # Looks for identical rows
          by=id]                         # (in var1 and var2 only!)

> head(res, 3)
   id    V1
1:  A  TRUE
2:  B FALSE
3:  C FALSE

一切似乎都有效,但是我的真实表(几乎80M行和4.5M unique(DT$id))我的代码需要2,1个小时。

有没有人有一些提示来加快上面的代码?我最终没有遵循从data.table功能中受益所需的最佳实践吗?提前谢谢任何人!

编辑:

将我的代码与@Arun&#39>进行比较的一些时间:

DT <- data.table(id   = rep(LETTERS,each=10000),
                 var1 = rnorm(260000),
                 var2 = rnorm(260000))

DT[, id2 := sample(c("A","B"), 10000, T), by=id] # I need this to simulate different 

setkey(DT)

> system.time(unique(DT)[, any(duplicated(.SD)), by = id, .SDcols = c("var1", "var2")])
   user  system elapsed 
   0.48    0.00    0.49 
> system.time(DT[, {a=unique(.SD)[,-3,with=F]   
+                   any(duplicated(a))}, 
+    by=id])
   user  system elapsed 
   1.09    0.00    1.10 

我想我得到了我想要的东西!

1 个答案:

答案 0 :(得分:4)

这个怎么样?

unique(setkey(DT))[, any(duplicated(.SD)), by=id, .SDcols = c("var1", "var2")]

在我的“慢速”机器上设置密钥大约需要140秒。实际的分组仍然在继续...... :)


这是我正在测试的巨大数据:

set.seed(1234)
DT <- data.table(id = rep(1:4500000, each=10), 
                 var1 = sample(1000, 45000000, replace=TRUE), 
                 var2 = sample(1000, 45000000, replace=TRUE))
DT[, id2 := sample(c("A","B"), 10, TRUE), by=id]