分区键部件网址

时间:2016-02-01 06:30:24

标签: apache-spark cassandra cassandra-2.0 cql3 spark-cassandra-connector

我有以下代码尝试在spark中加入2个cassandra表。

 val imageKeywords = sc.cassandraTable[ImageMetadata]("images", "metadata")
 val imageAndPageKeywords = imageKeywords
  .joinWithCassandraTable[PagesMetadata]("pages2", "metadata")
  .on(SomeColumns("tid", "url" as "pu"))

我用来映射数据的案例类如下

case class ImageMetadata(tid: String, iu: String, pu: Option[String],
mk: List[String], fk: List[String], ak: List[String], ipk: List[String], pk: List[String], ik: List[String], ck: List[String])

case class PagesMetadata(tid: String, url: String, pk: List[String], uk: List[String], hk: List[String], ok: List[String], tc: List[String])

当我尝试执行下面的操作时出现错误

imageAndPageKeywords.collect.toList.sortBy(_._1.tid).take(10).foreach(println)

错误堆栈跟踪 -

  

引起:com.datastax.driver.core.exceptions.InvalidQueryException:分区键部件url的null值无效       at com.datastax.driver.core.Responses $ Error.asException(Responses.java:103)       at com.datastax.driver.core.DefaultResultSetFuture.onSet(DefaultResultSetFuture.java:140)       在com.datastax.driver.core.RequestHandler.setFinalResult(RequestHandler.java:293)       在com.datastax.driver.core.RequestHandler.onSet(RequestHandler.java:455)       at com.datastax.driver.core.Connection $ Dispatcher.messageReceived(Connection.java:734)       在org.jboss.netty.channel.SimpleChannelUpstreamHandler.handleUpstream(SimpleChannelUpstreamHandler.java:70)       在org.jboss.netty.handler.timeout.IdleStateAwareChannelUpstreamHandler.handleUpstream(IdleStateAwareChannelUpstreamHandler.java:36)       在org.jboss.netty.channel.DefaultChannelPipeline.sendUpstream(DefaultChannelPipeline.java:564)       在org.jboss.netty.channel.DefaultChannelPipeline $ DefaultChannelHandlerContext.sendUpstream(DefaultChannelPipeline.java:791)       at org.jboss.netty.handler.timeout.IdleStateHandler.messageReceived(IdleStateHandler.java:294)       在org.jboss.netty.channel.SimpleChannelUpstreamHandler.handleUpstream(SimpleChannelUpstreamHandler.java:70)       在org.jboss.netty.channel.DefaultChannelPipeline.sendUpstream(DefaultChannelPipeline.java:564)       在org.jboss.netty.channel.DefaultChannelPipeline $ DefaultChannelHandlerContext.sendUpstream(DefaultChannelPipeline.java:791)       在org.jboss.netty.channel.Channels.fireMessageReceived(Channels.java:296)       在org.jboss.netty.handler.codec.oneone.OneToOneDecoder.handleUpstream(OneToOneDecoder.java:70)       在org.jboss.netty.channel.DefaultChannelPipeline.sendUpstream(DefaultChannelPipeline.java:564)       在org.jboss.netty.channel.DefaultChannelPipeline $ DefaultChannelHandlerContext.sendUpstream(DefaultChannelPipeline.java:791)       在org.jboss.netty.channel.Channels.fireMessageReceived(Channels.java:296)       在org.jboss.netty.handler.codec.frame.FrameDecoder.unfoldAndFireMessageReceived(FrameDecoder.java:462)       在org.jboss.netty.handler.codec.frame.FrameDecoder.callDecode(FrameDecoder.java:443)       在org.jboss.netty.handler.codec.frame.FrameDecoder.messageReceived(FrameDecoder.java:303)       在org.jboss.netty.channel.SimpleChannelUpstreamHandler.handleUpstream(SimpleChannelUpstreamHandler.java:70)       在org.jboss.netty.channel.DefaultChannelPipeline.sendUpstream(DefaultChannelPipeline.java:564)       在org.jboss.netty.channel.DefaultChannelPipeline.sendUpstream(DefaultChannelPipeline.java:559)       在org.jboss.netty.channel.Channels.fireMessageReceived(Channels.java:268)       在org.jboss.netty.channel.Channels.fireMessageReceived(Channels.java:255)       在org.jboss.netty.channel.socket.nio.NioWorker.read(NioWorker.java:88)       在org.jboss.netty.channel.socket.nio.AbstractNioWorker.process(AbstractNioWorker.java:108)       在org.jboss.netty.channel.socket.nio.AbstractNioSelector.run(AbstractNioSelector.java:318)       在org.jboss.netty.channel.socket.nio.AbstractNioWorker.run(AbstractNioWorker.java:89)       在org.jboss.netty.channel.socket.nio.NioWorker.run(NioWorker.java:178)       在org.jboss.netty.util.ThreadRenamingRunnable.run(ThreadRenamingRunnable.java:108)       在org.jboss.netty.util.internal.DeadLockProofWorker $ 1.run(DeadLockProofWorker.java:42)       ......还有3个

1 个答案:

答案 0 :(得分:2)

很简单,例外情况告诉您它无法执行连接,因为用于将 ImageMetadata PagesMetadata 连接的列为空。

在您的情况下, ImageMetadata 中的某些 url (pu)值为空。

奇怪的是你用 url nullable(Option [String])定义 PagesMetadata ,它似乎是表的主键的一部分

使其成功的一个解决方案是:

val imageAndPageKeywords = imageKeywords
  .filter(im -> im.pu.isDefined)
  .joinWithCassandraTable[PagesMetadata]("pages2", "metadata")
  .on(SomeColumns("tid", "url" as "pu"))