R中的连续值和新的因子水平

时间:2016-05-02 08:27:32

标签: r r-factor

我有以下示例

id <- c("a","b","a","b","a","a","a","a","b","b","c")
SOG <- c(4,4,0,0,0,0,0,0,0,0,9)
data <- data.frame(id,SOG)

我希望在新列中SOG == 0时的累积值。 使用以下代码

tmp <- rle(SOG)                                    #run length encoding: 
tmp$values <- tmp$values == 0                      #turn values into logicals 
tmp$values[tmp$values] <- cumsum(tmp$values[tmp$values]) #cumulative sum of TRUE values 
inverse.rle(tmp)                                   #inverse the run length encoding 

我创建了列&#34;停止&#34;:

data$Stops <- inverse.rle(tmp)

我可以进去:

[1] 0 0 1 1 1 1 1 1 1 1 0

但我想改为

[1] 0 0 1 2 3 3 3 3 4 4 0 

我的意思是当因素的水平&#34; id&#34;与前一行不同,我想跳到下一行&#34;停止&#34;第(i + 1)。

2 个答案:

答案 0 :(得分:4)

查看dplyr

library(dplyr)
data %>%
  mutate(
    Stops = ifelse(
      SOG > 0,
      0,
      cumsum(SOG == 0 & lag(id) != id)
    )
  )

答案 1 :(得分:1)

我们可以尝试

library(data.table)
setDT(data1)[, v1 := if(all(!SOG)) c(TRUE, id[-1]!= id[-.N]) else
     rep(FALSE, .N), .(grp = rleid(SOG))][,cumsum(v1)*(!SOG)]
#[1] 0 0 1 2 3 3 3 3 4 4 0 0 0 0 5 5 0 6 6 0

使用旧数据

setDT(data)[, v1 := if(all(!SOG)) c(TRUE, id[-1]!= id[-.N]) 
       else rep(FALSE, .N), .(grp = rleid(SOG))][,cumsum(v1)*(!SOG)]
#[1] 0 0 1 2 3 3 3 3 4 4 0

数据

id <- c("a","b","a","b","a","a","a","a","b","b","c","a","a","a","a","a","a","a","a", "a")
SOG <- c(4,4,0,0,0,0,0,0,0,0,9,1,5,3,0,0,4,0,0,1)
data1 <- data.frame(id, SOG, stringsAsFactors=FALSE)