三叉戟拓扑中的并行配置(风暴)

时间:2013-09-22 15:01:13

标签: apache-storm apache-kafka trident

阅读thisthis后,我无法理解如何配置三叉戟拓扑。

基本上我的风暴应用程序正在从kafka读取,进行一些数据操作并最终写入Cassandra

以下是我目前正在构建拓扑的方法:

private static StormTopology buildTopology() {
// connection to kafka
ZkHosts zkHosts = new ZkHosts(broker_zk, broker_path);
TridentKafkaConfig kafkaConfig = new TridentKafkaConfig(zkHosts, topic);
kafkaConfig.scheme = new RawMultiScheme();
StateFactoryFields[] cassandraStateFactories = createStateFactories();
TransactionalTridentKafkaSpout spout = new TransactionalTridentKafkaSpout(kafkaConfig);
TridentTopology topology = new TridentTopology();
Stream kafkaSpout = topology.newStream("kafkaspout", spout).parallelismHint(1).shuffle();
Stream filterValidatStream = kafkaSpout.each(new Fields("bytes"), new SplitKafkaInput(), EventData.getEventDataFields()).parallelismHint(1);
for (StateFactoryFields stateFactoryFields : cassandraStateFactories) {
    filterValidatStream.groupBy(stateFactoryFields.groupingFields)
        .persistentAggregate(stateFactoryFields.cassandraStateFactor, new Count(), new Fields("count")).parallelismHint(2);
}
logger.info("Building topology");
return topology.build();
}

所以我得到了一个spout和一些带有parallelismHint的操作(filter,grouopBy)。 我不明白确定最佳parallelismHint,而且如果我在我的代码中设置这个值,它如何与风暴标准拓扑配置一起工作,如

topology.max.task.parallelism
topology.workers
topology.acker.executors

提前致谢

1 个答案:

答案 0 :(得分:3)

有一个很好的gist by mrflip here试图概述如何调整风暴/三叉戟拓扑。这应该指导您选择参数(您在问题中建议的参数和您可能还没有想到的参数)。