Spark作业在本地运行时有效,但在独立模式下则无法工作

时间:2018-12-19 21:37:02

标签: java apache-spark apache-spark-sql

我有一个简单的Spark代码,在本地运行时可以正常工作,但是当我尝试将Spark Standalone Cluster与Docker一起运行时,它奇怪地失败了。

我可以确认与主人和工人的集成正在工作。

在下面的代码中,我显示了错误发生的地方。

JavaRDD<Row> rddwithoutMap = dataFrame.javaRDD();
JavaRDD<Row> rddwithMap = dataFrame.javaRDD()
            .map((Function<Row, Row>) row -> row);

long count = rddwithoutMap.count(); //here is fine
long countBeforeMap = rddwithMap.count(); // here I get the error

在地图之后,我无法调用任何Spark动作。

错误Caused by: java.lang.ClassNotFoundException: com.apssouza.lambda.MyApp$1

Obs:我在地图上使用Lambda,以使代码更具可读性,但在使用独立版本时,我也无法使用lambda。 Caused by: java.lang.ClassCastException: cannot assign instance of java.lang.invoke.SerializedLambda to field org.apache.spark.api.java.JavaPairRDD$$anonfun$toScalaFunction$1.fun$1 of type org.apache.spark.api.java.function.Function in instance of org.apache.spark.api.java.JavaPairRDD$$anonfun$toScalaFunction$1

Docker映像:bde2020/spark-master:2.3.2-hadoop2.7

本地Spark版本:2.4.0

火花依赖版本:spark-core_2.112.3.2

public class MyApp {
public static void main(String[] args) throws IOException, URISyntaxException {
//        String sparkMasterUrl = "local[*]";
//        String csvFile = "/Users/apssouza/Projetos/java/lambda-arch/data/spark/input/localhost.csv";

    String sparkMasterUrl = "spark://spark-master:7077";
    String csvFile = "hdfs://namenode:8020/user/lambda/localhost.csv";
    SparkConf sparkConf = new SparkConf()
            .setAppName("Lambda-demo")
            .setMaster(sparkMasterUrl);
         // .setJars(/path/to/my/jar); I even tried to set the jar
    JavaSparkContext sparkContext = new JavaSparkContext(sparkConf);
    SQLContext sqlContext = new SQLContext(sparkContext);
    Dataset<Row> dataFrame = sqlContext.read()
            .format("csv")
            .option("header", "true")
            .load(csvFile);

    JavaRDD<Row> rddwithoutMap = dataFrame.javaRDD();
    JavaRDD<Row> rddwithMap = dataFrame.javaRDD()
            .map((Function<Row, Row>) row -> row);

     long count = rddwithoutMap.count();
     long countBeforeMap = rddwithMap.count();

    }
}

<?xml version="1.0" encoding="UTF-8"?>

<project>
  <modelVersion>4.0.0</modelVersion>

  <groupId>com.apssouza.lambda</groupId>
  <artifactId>lambda-arch</artifactId>
  <version>1.0-SNAPSHOT</version>

  <name>lambda-arch</name>

  <properties>
    <project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
    <maven.compiler.source>1.8</maven.compiler.source>
    <maven.compiler.target>1.8</maven.compiler.target>
  </properties>

  <dependencies>

<dependency>
  <groupId>com.fasterxml.jackson.core</groupId>
  <artifactId>jackson-databind</artifactId>
  <version>2.9.7</version>
</dependency>

<dependency>
  <groupId>org.apache.spark</groupId>
  <artifactId>spark-core_2.11</artifactId>
  <version>2.3.2</version>
</dependency>
<!-- https://mvnrepository.com/artifact/org.apache.spark/spark-sql -->
<dependency>
  <groupId>org.apache.spark</groupId>
  <artifactId>spark-sql_2.11</artifactId>
  <version>2.3.2</version>
</dependency>

<dependency>
  <groupId>org.apache.commons</groupId>
  <artifactId>commons-lang3</artifactId>
  <version>3.6</version>
</dependency>

<dependency>
  <groupId>com.fasterxml.jackson.module</groupId>
  <artifactId>jackson-module-scala_2.11</artifactId>
  <version>2.9.7</version>
</dependency>

  </dependencies>


</project>

Obs:如果取消对前两行的注释,则一切正常。

1 个答案:

答案 0 :(得分:0)

问题是因为我在运行程序之前没有打包程序,并且在Spark集群中得到的应用程序版本过旧。这很奇怪,因为我正在通过我的IDE(IntelliJ)运行它,并且应该在运行它之前包装jar。无论如何,mvn package在点击运行按钮之前就解决了该问题。