Question

实木复合地板是由Spark v2.4 Parquet-mr v1.10生成的

n = 10000
x = [1.0, 2.0, 3.0, 4.0, 5.0, 5.0, None] * n
y = [u'é', u'é', u'é', u'é', u'a', None, u'a'] * n

z = np.random.rand(len(x)).tolist()
dfs = spark.createDataFrame(zip(x, y, z), schema=StructType([StructField('x', DoubleType(),True),StructField('y', StringType(), True),StructField('z', DoubleType(), False)]))
dfs.repartition(1).write.mode('overwrite').parquet('test_spark.parquet')

使用parquet-tools v1.12进行检查

row group 0 
--------------------------------------------------------------------------------
x:  DOUBLE SNAPPY DO:0 FPO:4 SZ:1632/31635/19.38 VC:70000 ENC:RLE,BIT_PACKED,PLAIN_DICTIONARY ST:[min: 1.0, max: 5.0, num_nulls: 10000]
y:  BINARY SNAPPY DO:0 FPO:1636 SZ:864/16573/19.18 VC:70000 ENC:RLE,BIT_PACKED,PLAIN_DICTIONARY ST:[min: a, max: é, num_nulls: 10000]
z:  DOUBLE SNAPPY DO:0 FPO:2500 SZ:560097/560067/1.00 VC:70000 ENC:PLAIN,BIT_PACKED ST:[min: 2.0828331581679294E-7, max: 0.9999892375625329, num_nulls: 0]

    x TV=70000 RL=0 DL=1 DS: 5 DE:PLAIN_DICTIONARY
    ----------------------------------------------------------------------------
    page 0:                   DLE:RLE RLE:BIT_PACKED VLE:PLAIN_DICTIONARY ST:[min: 1.0, max: 5.0, num_nulls: 10000] SZ:31514 VC:70000

    y TV=70000 RL=0 DL=1 DS: 2 DE:PLAIN_DICTIONARY
    ----------------------------------------------------------------------------
    page 0:                   DLE:RLE RLE:BIT_PACKED VLE:PLAIN_DICTIONARY ST:[min: a, max: é, num_nulls: 10000] SZ:16514 VC:70000

    z TV=70000 RL=0 DL=0
    ----------------------------------------------------------------------------
    page 0:                   DLE:BIT_PACKED RLE:BIT_PACKED VLE:PLAIN ST:[min: 2.0828331581679294E-7, max: 0.9999892375625329, num_nulls: 0] SZ:560000 VC:70000

问题：

FPO（第一数据页面偏移量）是否总是大于或小于DO（字典页面偏移量）？我从某个地方读取到字典页面存储在数据页面之后。

对于列x和y，plain_dictionary用于编码。但是，为什么两列的字典偏移都为0？

如果我使用使用parquet-cpp v1.5.1的pyarrow v0.11.1进行检查，它会告诉我has_dictionary_page: False和dictionary_page_offset: None

它是否有词典页面？

Answer 1

第一个数据页的偏移量始终大于词典的偏移量。换句话说，字典首先出现，然后才是数据页面。有两个用于存储这些偏移量的元数据字段：dictionary_page_offset（aka DO）和data_page_offset（aka FPO）。不幸的是，parquet-mr不能正确填写这些元数据字段。

例如，如果词典从偏移量1000开始，而第一个数据页从偏移量2000开始，则正确的值应为：

dictionary_page_offset = 1000
data_page_offset = 2000

相反，镶木地板先生商店

dictionary_page_offset = 0
data_page_offset = 1000

在您的示例中，这意味着尽管镶木地板工具显示了DO: 0，但x和y列仍然是字典编码的（z列不是）。

值得一提的是，Impala正确遵循了规范，因此您不能依赖每个具有此缺陷的文件。

这是parquet-mr在阅读过程中处理这种情况的方式：

// TODO: this should use getDictionaryPageOffset() but it isn't reliable.
if (f.getPos() != meta.getStartingPos()) {
  f.seek(meta.getStartingPos());
}

其中getStartingPos定义为：

/**
 * @return the offset of the first byte in the chunk
 */
public long getStartingPos() {
  long dictionaryPageOffset = getDictionaryPageOffset();
  long firstDataPageOffset = getFirstDataPageOffset();
  if (dictionaryPageOffset > 0 && dictionaryPageOffset < firstDataPageOffset) {
    // if there's a dictionary and it's before the first data page, start from there
    return dictionaryPageOffset;
  }
  return firstDataPageOffset;
}

您可以在上下文中看到以下几行代码：ParquetFileReader.readDictionary，ColumnChunkMetaData.getStartingPos。

为什么`plain_dictionary`编码的字典页面偏移量为0？

1 个答案: