我试图从数据框中取两个随机抽取的子样本,提取子样本中列的均值并计算均值之间的差异。以下函数和replicate
中do.call
的使用应该尽我所知,但我不断收到错误消息:
示例数据:
> dput(a)
structure(list(index = 1:30, val = c(14L, 22L, 1L, 25L, 3L, 34L,
35L, 36L, 24L, 35L, 33L, 31L, 30L, 30L, 29L, 28L, 26L, 12L, 41L,
36L, 32L, 37L, 56L, 34L, 23L, 24L, 28L, 22L, 10L, 19L), id = c(1L,
2L, 2L, 3L, 3L, 4L, 5L, 6L, 7L, 7L, 8L, 9L, 10L, 11L, 12L, 13L,
14L, 15L, 16L, 16L, 17L, 18L, 19L, 20L, 21L, 21L, 22L, 23L, 24L,
25L)), .Names = c("index", "val", "id"), class = "data.frame", row.names = c(NA,
-30L))
代码:
# Function to select only one row for each unique id in the data frame,
# take 2 randomly drawn subsets of size 40 from this unique dataset,
# calculate means of both subsets and determine the difference between the two means
extractDiff <- function(P){
xA <- ddply(P, .(id), function(x) {x[sample(nrow(x), 1) ,] }) # selects only one row for each id in the data frame
subA <- xA[sample(xA, 10, replace=TRUE), ] # takes a random sample of 40 rows
subB <- xA[sample(xA, 10, replace=TRUE), ] # takes a second random sample of 40 rows
meanA <- mean(subA$val)
meanB <- mean(subB$val)
diff <- abs(meanA-meanB)
outdf <- c(mA = meanA, mB= meanB, diffAB = diff)
return(outdf)
}
# To repeat the random selections and mean comparison X number of times...
fin <- do.call(rbind, replicate(10, extractDiff(a), simplify=FALSE))
错误讯息:
Error in xj[i] : invalid subscript type 'list'
我认为错误与不以可以提供给rbind
的格式返回函数输出有关,但我尝试的任何东西似乎都没有用(即我尝试将outdf对象转换为数据框和矩阵仍然得到错误moessage)。
我还在学习R所以对任何帮助都会感激不尽。谢谢!
答案 0 :(得分:0)
如果您将sample
列表/ data.frame作为第一个参数传递,它将返回list / data.frame。您无法使用data.frame对data.frame进行子集化。
library(plyr)
extractDiff <- function(P){
xA <- ddply(P, .(id), function(x) {x[sample(nrow(x), 1) ,] }) # selects only one row for each id in the data frame
subA <- xA[sample(nrow(xA), 10, replace=TRUE), ] # takes a random sample of 40 rows
subB <- xA[sample(nrow(xA), 10, replace=TRUE), ] # takes a second random sample of 40 rows
meanA <- mean(subA$val)
meanB <- mean(subB$val)
diff <- abs(meanA-meanB)
outdf <- c(mA = meanA, mB= meanB, diffAB = diff)
return(outdf)
}
set.seed(42)
fin <- do.call(rbind, replicate(10, extractDiff(a), simplify=FALSE))
# mA mB diffAB
# [1,] 29.4 25.5 3.9
# [2,] 25.8 23.0 2.8
# [3,] 25.3 29.5 4.2
# [4,] 29.0 31.2 2.2
# [5,] 26.5 25.6 0.9
# [6,] 26.8 27.2 0.4
# [7,] 28.7 27.3 1.4
# [8,] 22.7 28.7 6.0
# [9,] 30.6 23.2 7.4
# [10,] 25.1 25.2 0.1