如何考虑文件名来重命名文件名

时间:2019-01-23 21:17:59

标签: r data-analysis data-manipulation

我是R编程的入门者。我已经下载了许多以ID为名称的图片。例如,图片“ senador588”,“ senador3”,“ senador16”等等。每张照片都显示了一位巴西参议员。我需要名称而不是ID。

我还有一个数据框,仅显示ID(id_senador)和名称(name_lower)。

代码的第一部分下载了所有图片:

Remote server did not respond

第二部分创建一个具有每个参议员的ID和名称的数据框:

library(data.table)
library(rvest)
library(lubridate)
library(stringr)
library(dplyr)
library(RCurl)
library(XML)
library(httr)
library(purrr)
# all the senators of Brazil
url <- "https://www25.senado.leg.br/web/senadores/em-exercicio/-/e/por-nome"


# get all url on the webpage
url2 <- getURL(url)
parsed <- htmlParse(url2)
links <- xpathSApply(parsed,path = "//a",xmlGetAttr,"href")

links <- do.call(rbind.data.frame, links) 

colnames(links)[1] <- "links" 


# filtering to get the urls of the senators
links_senador <- links %>%
  filter(links %like% "/senadores/senador/")

links_senador <- data.frame(links_senador)

# creating a new directory for the pics
setwd("~/Downloads/")
dir.create("senadores-new")
setwd("~/Downloads/senadores-new")

# running a loop to download all pictures
i <- 1
while(1 <= 81){
  tryCatch({
# defining the row of each senator
  foto_webpage <- data.frame(links_senador$links[i])
# renaming the column's name
  colnames(foto_webpage) <- "links" 
# getting all images of html page
# filtering the photo which we want
  html <- as.character(foto_webpage$links) %>%
    httr::GET() %>%
    xml2::read_html() %>%
    rvest::html_nodes("img") %>%
    map(xml_attrs) %>%
    map_df(~as.list(.)) %>%
    filter(src %like% "senadores/img/fotos-oficiais/") %>%
    as.data.frame(html)
# downloading the photo
    foto_senador <- html$src
    download.file(foto_senador, basename(foto_senador), mode = "wb", header = TRUE)
    Sys.sleep(3)
  }, error = function(e) return(NULL)
  )
  i <- i + 1
}

为了用循环替换名称的ID,我尝试了以下方法:

url <- "https://www25.senado.leg.br/web/senadores/em-exercicio/-/e/por-nome"

file <- read_html(url)
tables <- html_nodes(file, "table")
table1 <- html_table(tables[1], fill = TRUE, header = T)


table1_df <- as.data.frame(table1)[1]

table1_df_sem_acentuacao <- as.data.frame(iconv(table1_df$Nome, from = "UTF-8", to = "ASCII//TRANSLIT"))
colnames(table1_df_sem_acentuacao) <- "senador_lower"

table1_df_lower <- as.data.frame(tolower(table1_df_sem_acentuacao$senador_lower))
colnames(table1_df_lower) <- "senador_lower"

table_name_final <- as.data.frame(gsub(" ", "-", table1_df_lower$senador_lower))

id_split <- as.data.frame(gsub("https://www25.senado.leg.br/web/senadores/senador/-/perfil/", "senador", links_senador$links))

table_dfs_final <- cbind(table_name_final, id_split)
colnames(table_dfs_final)[1] <- "name_lower"
colnames(table_dfs_final)[2] <- "id_senador"

1 个答案:

答案 0 :(得分:1)

要使其更具“ R方式”,可以使用apply系列的功能之一。创建可以更改名称的函数,而不仅仅是将其应用于您创建的id和名称列。

for