更改镶木地板文件中的列数据类型

时间:2018-03-31 12:02:10

标签: sql amazon-web-services amazon-s3 hive external-tables

我有一个指向s3位置(镶木地板文件)的外部表,其中所有数据类型都是字符串。我想纠正所有列的数据类型,而不是只读取所有列的字符串。当我删除外部表并使用新的数据类型重新创建时,select查询总是抛出错误,如下所示:

java.lang.UnsupportedOperationException: org.apache.parquet.column.values.dictionary.PlainValuesDictionary$PlainBinaryDictionary
    at org.apache.parquet.column.Dictionary.decodeToInt(Dictionary.java:48)
    at org.apache.spark.sql.execution.vectorized.OnHeapColumnVector.getInt(OnHeapColumnVector.java:233)
    at org.apache.spark.sql.catalyst.expressions.GeneratedClass$GeneratedIterator.processNext(Unknown Source)
    at org.apache.spark.sql.execution.BufferedRowIterator.hasNext(BufferedRowIterator.java:43)
    at org.apache.spark.sql.execution.WholeStageCodegenExec$$anonfun$8$$anon$1.hasNext(WholeStageCodegenExec.scala:395)
    at org.apache.spark.sql.execution.SparkPlan$$anonfun$2.apply(SparkPlan.scala:234)
    at org.apache.spark.sql.execution.SparkPlan$$anonfun$2.apply(SparkPlan.scala:228)
    at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$25.apply(RDD.scala:827)
    at org.apache.spark.rdd.RDD$$anonfun$mapPartitionsInternal$1$$anonfun$apply$25.apply(RDD.scala:827)

1 个答案:

答案 0 :(得分:0)

将类型指定为BigInt,它等效于long类型,hive没有long数据类型。

hive> alter table table change col col bigint;
  

来自Hortonworks论坛的重复内容