使用emr中的spark从S3中读取avro失败

时间:2017-10-16 14:36:13

标签: hadoop apache-spark apache-spark-sql amazon-emr spark-avro

在aws-emr上执行我的Spark作业时,在尝试从s3存储桶读取avro文件时出现此错误: 它发生在版本中:

  • emr - 5.5.0
  • emr - 5.9.0

这是代码:

val files  = 0 until numOfDaysToFetch map { i =>
  s"s3n://bravos/clicks/${fromDate.minusDays(i)}/*"
}
spark.read.format("com.databricks.spark.avro").load(files: _*)

例外:

java.lang.IllegalArgumentException: java.net.URISyntaxException: Relative path in absolute URI: 1037330823653531755-2017-10-16T03:06:00.avro
    at org.apache.hadoop.fs.Path.initialize(Path.java:205)
    at org.apache.hadoop.fs.Path.<init>(Path.java:171)
    at org.apache.hadoop.fs.Path.<init>(Path.java:93)
    at org.apache.hadoop.fs.Globber.glob(Globber.java:241)
    at org.apache.hadoop.fs.FileSystem.globStatus(FileSystem.java:1732)
    at org.apache.hadoop.fs.FileSystem.globStatus(FileSystem.java:1713)
    at com.amazon.ws.emr.hadoop.fs.EmrFileSystem.globStatus(EmrFileSystem.java:362)
    at org.apache.spark.deploy.SparkHadoopUtil.globPath(SparkHadoopUtil.scala:237)
    at org.apache.spark.deploy.SparkHadoopUtil.globPathIfNecessary(SparkHadoopUtil.scala:243)
    at org.apache.spark.sql.execution.datasources.DataSource$$anonfun$14.apply(DataSource.scala:374)
    at org.apache.spark.sql.execution.datasources.DataSource$$anonfun$14.apply(DataSource.scala:370)
    at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)
    at scala.collection.TraversableLike$$anonfun$flatMap$1.apply(TraversableLike.scala:241)
    at scala.collection.immutable.List.foreach(List.scala:381)
    at scala.collection.TraversableLike$class.flatMap(TraversableLike.scala:241)
    at scala.collection.immutable.List.flatMap(List.scala:344)
    at org.apache.spark.sql.execution.datasources.DataSource.resolveRelation(DataSource.scala:370)
    at org.apache.spark.sql.DataFrameReader.load(DataFrameReader.scala:152)

`

2 个答案:

答案 0 :(得分:0)

Path不支持冒号。它解释1037330823653531755-2017-10-16T03:作为一个URI模式,然后对任何填充&#34; /&#34; ...感到不满意。即使它到达那么远,它也会失败#34 ;没有架构的文件系统&#34; 1037330823653531755-2017-10-16T03&#34;

修正:不要使用&#34;:&#34;在文件名中。

答案 1 :(得分:0)

我从/ *删除了最后一个*它刚刚工作