
时间:2019-01-04 14:53:14

标签: r tar tidyverse



sort -u

仅使用R 是否可以直接将其中一个文件读/加载到R中(又不需要先将文件解压缩并将其写入磁盘)?

1 个答案:

答案 0 :(得分:1)


第一个问题是,您应该仅从头开始读取压缩文件。不幸的是,在gzip压缩的文件中,使用“ seek()”将文件指针重新定位到所需文件的位置不稳定。

ParseTGZ<- function(archname){
  # open tgz archive
  tf <- gzfile(archname, open='rb')
  fnames <- list()
  offset <- 0
  nfile <- 0
  while (TRUE) {
    # go to beginning of entry
    # never use "seek" to re-locate in a gzipped file!
    if (seek(tf) != offset) readBin(tf, what="raw", n= offset - seek(tf))
    # read file name
    fName <- rawToChar(readBin(tf, what="raw", n=100))
    if (nchar(fName)==0) break
    nfile <- nfile + 1
    fnames <- c(fnames, fName)
    attr(fnames[[nfile]], "offset") <- offset+512
    # read size, first skip 24 bytes (file permissions etc)
    # again, we only use readBin, not seek()
    readBin(tf, what="raw", n=24)
    # file size is encoded as a length 12 octal string, 
    # with the last character being '\0' (so 11 actual characters)
    sz <- readChar(tf, nchars=11) 
    # convert string to number of bytes
    sz <- sum(as.numeric(strsplit(sz,'')[[1]])*8^(10:0))
    attr(fnames[[nfile]], "size") <- sz
#    cat(sprintf('entry %s, %i bytes\n', fName, sz))
    # go to the next message
    # don't forget entry header (=512) 
    offset <- offset + 512*(ceiling(sz/512) + 1)
# return a named list of characters strings with attributes?
  names(fnames) <- fnames

这将为您提供tar.gz归档文件中所有文件的确切位置和长度。 现在,下一步是实际扩展单个文件。您可能可以直接使用“ gzfile”连接来执行此操作,但是在这里我将使用rawConnection()。假定您的文件适合内存。

extractTGZ <- function(archfile, filename) {
  # this function returns a raw vector
  # containing the desired file
  fp <- ParseTGZ(archfile)
  offset <- attributes(fp[[filename]])$offset
  fsize <- attributes(fp[[filename]])$size
  gzf <- gzfile(archfile, open="rb")
  # jump to the byte position, don't use seek()
  # may be a bad idea on really large archives...
  readBin(gzf, what="raw", n=offset)
  # now read the data into a raw vector
  result <- readBin(gzf, what="raw", n=fsize)


ff <- rawConnection(ExtractTGZ("myarchive", "myfile"))
