R:将可变数量的行粘贴(或合并)为一个

时间:2018-10-30 15:36:30

标签: r parsing paste

我有一个试图解析的文本文件,并将信息放入数据框。在每个“事件”中,可能有也可能没有注释。但是,注释可以跨越各种行。我需要将每个事件的注释连接为一个字符串,以存储在数据框的一列中。

ID: 20470
Version: 1
notes: 


ID: 01040
Version: 2
notes: 
The customer was late.
Project took 20 min. longer than anticipated
Work was successfully completed

ID: 00000
Version: 1
notes: 
Customer was not at home.

ID: 00000
Version: 7
notes: 
Fax at 2:30 pm
Called but no answer
Visit home no answer
Left note on door with call back number
Made a final attempt on 12/5/2013
closed case on 12/10 with nothing resolved 

例如,对于第三次事件,注释应为一个长字符串:“客户迟到。项目花费了比预期工作成功完成要长20分钟的时间”,然后将其存储到数据框。

对于每个事件,我知道注释跨越多少行。

1 个答案:

答案 0 :(得分:0)

这样的事情(实际上,您会更快乐,并且可以自己了解更多信息,我只是在两个任务之间拖延时间):

x <- readLines("R/xample.txt")  # you'll probably read it from  a file
ids <- grep("^ID:", x)   # detecting lines starting with ID:
versions <- grep("^Version:", x)
notes <- grep("^notes:", x)
nStart <- notes + 1  # lines where the notes start
nEnd <- c(ids[-1]-1, length(x))  # notes end one line before the next ID: line
ids <- sapply(strsplit(x[ids], ": "), "[[", 2)
versions <- sapply(strsplit(x[versions], ": "), "[[", 2)
notes <- mapply(function(i,j) paste(x[i:j], collapse=" "), nStart, nEnd)
df <- data.frame(ID=ids, ver=versions, note=notes, stringsAsFactors=FALSE)

数据输入

> dput(x)
c("ID: 20470", "Version: 1", "notes: ", "  ", "  ", "ID: 01040", 
"Version: 2", "notes: ", "  The customer was late.", "Project took 20 min. longer than anticipated", 
"Work was successfully completed", "", "ID: 00000", "Version: 1", 
"notes: ", "  Customer was not at home.", "", "ID: 00000", "Version: 7", 
"notes: ", "  Fax at 2:30 pm", "Called but no answer", "Visit home no answer", 
"Left note on door with call back number", "Made a final attempt on 12/5/2013", 
"closed case on 12/10 with nothing resolved ")