在Apache Solr云中索引巨大的表记录

时间:2018-12-28 16:31:31

标签: apache solr cassandra

我有一个Cassandra表,其中有900万条记录,数据大小为500 MB。我有一个带有3个节点(3个碎片和2个副本)和3个外部Zookeeper集成的Solr云。我的Cassandra是1节点群集。我正在尝试使用Apache Solr为该表建立索引,但是一旦我开始完全导入,我的查询就会超时。

我能够cqlsh并获取记录,但是无法为其编制索引。 这是我所附的solr.log ...

Caused by: org.apache.solr.handler.dataimport.DataImportHandlerException: Unable to execute query: SELECT * from counter.series Processing Document # 1
        at org.apache.solr.handler.dataimport.DataImportHandlerException.wrapAndThrow(DataImportHandlerException.java:69)
        at org.apache.solr.handler.dataimport.JdbcDataSource$ResultSetIterator.<init>(JdbcDataSource.java:318)
        at org.apache.solr.handler.dataimport.JdbcDataSource.getData(JdbcDataSource.java:279)
        at org.apache.solr.handler.dataimport.JdbcDataSource.getData(JdbcDataSource.java:54)
        at org.apache.solr.handler.dataimport.SqlEntityProcessor.initQuery(SqlEntityProcessor.java:59)
        at org.apache.solr.handler.dataimport.SqlEntityProcessor.nextRow(SqlEntityProcessor.java:73)
        at org.apache.solr.handler.dataimport.EntityProcessorWrapper.nextRow(EntityProcessorWrapper.java:244)
        at org.apache.solr.handler.dataimport.DocBuilder.buildDocument(DocBuilder.java:475)
        at org.apache.solr.handler.dataimport.DocBuilder.buildDocument(DocBuilder.java:414)
        ... 5 more
Caused by: java.sql.SQLTransientConnectionException: TimedOutException()
        at org.apache.cassandra.cql.jdbc.CassandraStatement.doExecute(CassandraStatement.java:189)
        at org.apache.cassandra.cql.jdbc.CassandraStatement.execute(CassandraStatement.java:205)
        at org.apache.solr.handler.dataimport.JdbcDataSource$ResultSetIterator.executeStatement(JdbcDataSource.java:338)
        at org.apache.solr.handler.dataimport.JdbcDataSource$ResultSetIterator.<init>(JdbcDataSource.java:313)
        ... 12 more
Caused by: TimedOutException()
        at org.apache.cassandra.thrift.Cassandra$execute_cql3_query_result.read(Cassandra.java:37865)
        at org.apache.thrift.TServiceClient.receiveBase(TServiceClient.java:78)
        at org.apache.cassandra.thrift.Cassandra$Client.recv_execute_cql3_query(Cassandra.java:1562)
        at org.apache.cassandra.thrift.Cassandra$Client.execute_cql3_query(Cassandra.java:1547)
        at org.apache.cassandra.cql.jdbc.CassandraConnection.execute(CassandraConnection.java:468)
        at org.apache.cassandra.cql.jdbc.CassandraConnection.execute(CassandraConnection.java:494)
        at org.apache.cassandra.cql.jdbc.CassandraStatement.doExecute(CassandraStatement.java:164)
        ... 15 more

我希望以批处理或通过使用多个线程的方式为表建立索引。欢迎任何帮助或建议。 db-data-config.xml: <dataConfig> <dataSource type="JdbcDataSource" driver="org.apache.cassandra.cql.jdbc.CassandraDriver" url="jdbc:cassandra://192.168.0.7:9160/counter" user="cassandra" password="cassandra" autoCommit="true" /> <document> <entity name="counter" query="SELECT * from counter.series;" autoCommit="true"> <field column="serial" name="serial" /> <field column="random" name="random" /> <field column="remarks" name="remarks" /> <field column="timestamp" name="timestamp" /> </entity> </document> </dataConfig>

solrconfig.xml

`<requestHandler name="/dataimport" class="org.apache.solr.handler.dataimport.DataImportHandler">
<lst name="defaults">
  <str name="config">db-data-config.xml</str>
</lst>

`

schema.xml <field name="remarks" type="string" indexed="false" stored="false" required="false" /> <field name="serial" type="string" indexed="true" stored="true" required="true" /> <field name="random" type="string" indexed="false" stored="true" required="true" /> <field name="timestamp" type="string" indexed="false" stored="false" required="false" />

1 个答案:

答案 0 :(得分:0)

问题很可能是发送到Solr的数据有效负载的大小。默认情况下,当batchSize中未指定任何JdbcDataSource时,默认值为500。在您的情况下,它看起来太多了。您应该在Solr端使用较小的数字或增加超时设置