我有两个看起来像这样的数据帧(虽然第一个数据帧超过9000万行,第二个数据帧有超过1400万行)另外第二个数据帧是随机排序的
df1 <- data.frame(
datalist = c("wiki/anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/individualism to complete wiki/collectivism",
"strains of anarchism have often been divided into the categories of wiki/social_anarchism and wiki/individualist_anarchism or similar dual classifications",
"the word is composed from the word wiki/anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e",
"anarchy from anarchos meaning one without rulers from the wiki/privative prefix wiki/privative_alpha an- i.e",
"authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/infinitive suffix -izein",
"the first known use of this word was in 1539"),
words = c("anarchist_schools_of_thought individualism collectivism", "social_anarchism individualist_anarchism",
"anarchy -ism", "privative privative_alpha", "infinitive", ""),
stringsAsFactors=FALSE)
df2 <- data.frame(
vocabword = c("anarchist_schools_of_thought", "individualism","collectivism" , "1965-66_nhl_season_by_team","social_anarchism","individualist_anarchism",
"anarchy","-ism","privative","privative_alpha", "1310_the_ticket", "infinitive"),
token = c("Anarchist_schools_of_thought" ,"Individualism", "Collectivism", "1965-66_NHL_season_by_team", "Social_anarchism", "Individualist_anarchism" ,"Anarchy",
"-ism", "Privative" ,"Alpha_privative", "KTCK_(AM)" ,"Infinitive"),
stringsAsFactors = F)
我能够将短语“wiki /”之后的所有单词提取到另一列中。这些单词需要由与第二个数据帧中的vocabword匹配的标记列替换。因此,举例来说,我会看一下wiki /在第一个数据帧的第一行之后的“anarchist_schools_of_thought”作品,然后在词汇表下的第二个数据框中找到术语“anarchist_schools_of_thought”,我想用相应的替换它令牌是“Anarchist_schools_of_thought”。
所以最终应该看起来像这样:
1 wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism
2 strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications
3 the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e
4 anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_privative an- i.e
5 authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein
6 the first known use of this word was in 1539
我意识到他们中的很多人只是将这些单词的第一个字母大写,但其中一些字母明显不同。我可以做一个for循环,但我认为这会花费太多时间,我更喜欢以data.table方式或者可能是stringi或stringr方式。我通常只会进行合并,但由于需要在一行中替换多个单词,因此会使事情复杂化。
提前感谢您的帮助。
答案 0 :(得分:1)
您可以使用str_replace_all
中的stringr
:
library(stringr)
str_replace_all(df1$datalist, setNames(df2$vocabword, df2$token))
基本上,str_replace_all
允许您提供一个命名向量,其中原始字符串是名称,替换是向量的元素。你通过创建一个&#34;字典来完成所有艰苦的工作。字符串和替换。 str_replace_all
只是简单地接受并自动进行替换。
<强>结果:强>
[1] "wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism"
[2] "strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications"
[3] "the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e"
[4] "Anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Privative_alpha an- i.e"
[5] "authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein"
[6] "the first known use of this word was in 1539"
答案 1 :(得分:0)
此问题的解决方案似乎与您的数据配合良好:R: replacing multiple regex with sub
install.packages('qdap')
qdap::mgsub(df2[,1], df2[,2], df1[,1])
[1] "wiki/Anarchist_schools_of_thought can differ fundamentally supporting anything from extreme wiki/Individualism to complete wiki/Collectivism"
[2] "strains of anarchism have often been divided into the categories of wiki/Social_anarchism and wiki/Individualist_anarchism or similar dual classifications"
[3] "the word is composed from the word wiki/Anarchy and the suffix wiki/-ism themselves derived respectively from the greek i.e"
[4] "Anarchy from anarchos meaning one without rulers from the wiki/Privative prefix wiki/Alpha_Privative an- i.e"
[5] "authority sovereignty realm magistracy and the suffix or -ismos -isma from the verbal wiki/Infinitive suffix -izein"
[6] "the first known use of this word was in 1539"