我在亚马逊上有一个由Priam管理的Cassandra集群上运行的Java应用程序。
我们使用亚马逊的Elastic Map / Reduce服务,在某些时刻,当我运行EMR并尝试在Cassandra上插入一些数据时,我得到了一个异常:OperationTimeoutException。
这些是我在Astyanax上创建Cassandra池时传递的配置参数:
`ConnectionPoolConfigurationImpl conPool = new` `ConnectionPoolConfigurationImpl(getConecPoolName())`
.setMaxConnsPerHost(20)
.setSeeds("ec2-xx-xxx-xx-xx.compute-1.amazonaws.com")
.setMaxOperationsPerConnection(100) .setMaxPendingConnectionsPerHost(20)
.setConnectionLimiterMaxPendingCount(20)
.setTimeoutWindow(10000)
.setConnectionLimiterWindowSize(1000)
.setMaxTimeoutCount(3)
.setConnectTimeout(5000)
.setMaxFailoverCount(-1)
.setLatencyAwareBadnessThreshold(20)
.setLatencyAwareUpdateInterval(1000)
.setLatencyAwareResetInterval(10000)
.setLatencyAwareWindowSize(100)
.setLatencyAwareSentinelCompare(100f)
AstyanaxContext<Keyspace> context = new AstyanaxContext.Builder()
.forCluster("clusterName")
.forKeyspace("keyspaceName")
.withAstyanaxConfiguration(
new AstyanaxConfigurationImpl().setDiscoveryType(NodeDiscoveryType.NONE))
.withConnectionPoolConfiguration(conPool)
.withConnectionPoolMonitor(new CountingConnectionPoolMonitor())
.buildKeyspace(ThriftFamilyFactory.getInstance());
完整堆栈跟踪:
ERROR com.s1mbi0se.dg.input.service.InputService (main): EXCEPTION:OperationTimeoutException: [host=ec2-xx-xxx-xx-xx.compute-1.amazonaws.com(10.100.6.242):9160, latency=10004(10004), attempts=1]TimedOutException()
com.netflix.astyanax.connectionpool.exceptions.OperationTimeoutException: OperationTimeoutException: [host=ec2-54-224-65-18.compute-1.amazonaws.com(10.100.6.242):9160, latency=10004(10004), attempts=1]TimedOutException()
at com.netflix.astyanax.thrift.ThriftConverter.ToConnectionPoolException(ThriftConverter.java:171)
at com.netflix.astyanax.thrift.AbstractOperationImpl.execute(AbstractOperationImpl.java:61)
at com.netflix.astyanax.thrift.ThriftColumnFamilyQueryImpl$1$2.execute(ThriftColumnFamilyQueryImpl.java:206)
at com.netflix.astyanax.thrift.ThriftColumnFamilyQueryImpl$1$2.execute(ThriftColumnFamilyQueryImpl.java:198)
at com.netflix.astyanax.thrift.ThriftSyncConnectionFactoryImpl$ThriftConnection.execute(ThriftSyncConnectionFactoryImpl.java:151)
at com.netflix.astyanax.connectionpool.impl.AbstractExecuteWithFailoverImpl.tryOperation(AbstractExecuteWithFailoverImpl.java:69)
at com.netflix.astyanax.connectionpool.impl.AbstractHostPartitionConnectionPool.executeWithFailover(AbstractHostPartitionConnectionPool.java:253)
at com.netflix.astyanax.thrift.ThriftColumnFamilyQueryImpl$1.execute(ThriftColumnFamilyQueryImpl.java:196)
at com.s1mbi0se.dg.input.service.InputService.searchUserByKey(InputService.java:833)
at org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:144)
at org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:771)
at org.apache.hadoop.mapred.MapTask.run(MapTask.java:375)
at org.apache.hadoop.mapred.Child$4.run(Child.java:255)
at java.security.AccessController.doPrivileged(Native Method)
at javax.security.auth.Subject.doAs(Subject.java:396)
at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1132)
at org.apache.hadoop.mapred.Child.main(Child.java:249)
Caused by: TimedOutException()
at org.apache.cassandra.thrift.Cassandra$get_slice_result.read(Cassandra.java:7874)
at org.apache.thrift.TServiceClient.receiveBase(TServiceClient.java:78)
at org.apache.cassandra.thrift.Cassandra$Client.recv_get_slice(Cassandra.java:594)
at org.apache.cassandra.thrift.Cassandra$Client.get_slice(Cassandra.java:578)
at com.netflix.astyanax.thrift.ThriftColumnFamilyQueryImpl$1$2.internalExecute(ThriftColumnFamilyQueryImpl.java:211)
at com.netflix.astyanax.thrift.ThriftColumnFamilyQueryImpl$1$2.internalExecute(ThriftColumnFamilyQueryImpl.java:198)
at com.netflix.astyanax.thrift.AbstractOperationImpl.execute(AbstractOperationImpl.java:56)
所以我不知道我要解决这个问题的方向,因为问题可能出在Astyanax池配置,EC2机器配置(内存增加?),Priam配置或Cassandra或EMR所需的其他配置上我的代码中的AWS服务...任何提示?
跟随堆栈跟踪:
INFO org.apache.hadoop.mapred.TaskLogsTruncater (main): Initializing logs' truncater with mapRetainSize=-1 and reduceRetainSize=-1
WARN org.apache.hadoop.mapred.Child (main): Error running child
java.lang.RuntimeException: InvalidRequestException(why:Start key's token sorts after end token)
at org.apache.cassandra.hadoop.ColumnFamilyRecordReader$WideRowIterator.maybeInit(ColumnFamilyRecordReader.java:453)
at org.apache.cassandra.hadoop.ColumnFamilyRecordReader$WideRowIterator.computeNext(ColumnFamilyRecordReader.java:459)
at org.apache.cassandra.hadoop.ColumnFamilyRecordReader$WideRowIterator.computeNext(ColumnFamilyRecordReader.java:406)
at com.google.common.collect.AbstractIterator.tryToComputeNext(AbstractIterator.java:143)
at com.google.common.collect.AbstractIterator.hasNext(AbstractIterator.java:138)
at org.apache.cassandra.hadoop.ColumnFamilyRecordReader.getProgress(ColumnFamilyRecordReader.java:103)
at org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.getProgress(MapTask.java:522)
at org.apache.hadoop.mapred.MapTask$NewTrackingRecordReader.nextKeyValue(MapTask.java:547)
at org.apache.hadoop.mapreduce.MapContext.nextKeyValue(MapContext.java:67)
at org.apache.hadoop.mapreduce.Mapper.run(Mapper.java:143)
at org.apache.hadoop.mapred.MapTask.runNewMapper(MapTask.java:771)
at org.apache.hadoop.mapred.MapTask.run(MapTask.java:375)
at org.apache.hadoop.mapred.Child$4.run(Child.java:255)
at java.security.AccessController.doPrivileged(Native Method)
at javax.security.auth.Subject.doAs(Subject.java:396)
at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1132)
at org.apache.hadoop.mapred.Child.main(Child.java:249)
Caused by: InvalidRequestException(why:Start key's token sorts after end token)
at org.apache.cassandra.thrift.Cassandra$get_paged_slice_result.read(Cassandra.java:14168)
at org.apache.thrift.TServiceClient.receiveBase(TServiceClient.java:78)
at org.apache.cassandra.thrift.Cassandra$Client.recv_get_paged_slice(Cassandra.java:769)
at org.apache.cassandra.thrift.Cassandra$Client.get_paged_slice(Cassandra.java:753)
at org.apache.cassandra.hadoop.ColumnFamilyRecordReader$WideRowIterator.maybeInit(ColumnFamilyRecordReader.java:438)
...
INFO org.apache.hadoop.mapred.Task (main): Runnning cleanup for the task
答案 0 :(得分:1)
我们解决了这个问题(Dean我在Cassandra用户组上回答了这个问题,但我会再次提出我们在这里解决问题的方法)
基本上我认为我们最初的问题是亚马逊延迟很高,所以我们改变了我们的池配置,而且工作正常......
感谢大家的帮助(主要是Dean)!
低于我们在亚马逊上运行的实际池配置:
new ConnectionPoolConfigurationImpl(getConecPoolName())
.setMaxConnsPerHost(CONNECTION_POOL_SIZE_PER_HOST)
.setSeeds(getIpSeeds())
.setMaxOperationsPerConnection(10000)
.setMaxPendingConnectionsPerHost(20)
.setConnectionLimiterMaxPendingCount(20)
.setTimeoutWindow(10000)
.setConnectionLimiterWindowSize(2000)
.setMaxTimeoutCount(3)
.setConnectTimeout(100)
.setConnectTimeout(2000)
.setMaxFailoverCount(-1)
.setLatencyAwareBadnessThreshold(20)
.setLatencyAwareUpdateInterval(1000) // 10000
.setLatencyAwareResetInterval(10000) // 60000
.setLatencyAwareWindowSize(100) // 100
.setLatencyAwareSentinelCompare(100f) .setSocketTimeout(30000)
.setMaxTimeoutWhenExhausted(10000)
.setInitConnsPerHost(10)
;
AstyanaxContext<Keyspace> context = new AstyanaxContext.Builder().forCluster(clusterName).forKeyspace(keyspaceName)
.withAstyanaxConfiguration(new AstyanaxConfigurationImpl().setDiscoveryType(NodeDiscoveryType.NONE).setConnectionPoolType(ConnectionPoolType.ROUND_ROBIN).setDiscoveryDelayInSeconds(10000)
.setDiscoveryDelayInSeconds(10000))
.withConnectionPoolConfiguration(conPool)
.withConnectionPoolMonitor(new CountingConnectionPoolMonitor())
.buildKeyspace(ThriftFamilyFactory.getInstance());
答案 1 :(得分:0)
那么,如果将超时设置为-1会发生什么?就个人而言,我会深入研究astyanax代码并尝试弄清楚如何禁用超时。再次运行你的东西它应该继续,虽然你的群集可能会受到打击,如果你得到超时...我假设你没关系。
编辑(编辑后):拍摄,我忘了问你正在使用哪个版本的cassandra。我正在看这个代码但是第346行是你的第438行(你使用的是宽度迭代器,这很可能意味着行扫描(即行的各个部分)。我们至少可以看到这是获取一个键范围,但是因为行可能太宽而被分页(对于内存不太宽的行,还有另一个迭代器)。我相信你是正确的,你不能使用两种分区类型。为了获得更多信息,我强烈建议修改ColumnFamilyRecordReader.java以记录ColumnFamilySplit(它上面有一个toString)。你可以在initialize方法中记录它,也可以记录JobRange(它也有一个toString)。
即
logger.warn("my split range="+split+" job's total range="+jobRange);
您的版本与此代码有很多相似之处。你在哪个版本?
除了拆分之外,我还会记录KeySlice以防万一,如果我没记错的话可能会导致错误。让我知道您的版本,并添加一些日志以获取有关您的情况的更多信息。 (他们的东西通常很容易开箱即用,没有任何问题)。
迪安