如何使用Hibernate Lucene Search进行不区分大小写的排序?

时间:2016-07-20 04:10:43

标签: java hibernate lucene hibernate-search

我可以使用以下代码获得结果,但结果未正确排序。它首先显示小写字母,然后显示大写字符。

获得结果:

upper
test
UPPER
Test

预期结果;

 upper
 UPPER
 Test
 test 

模式可以是大写(T)的第一个和小写(T)。

以下是供参考的代码:

普拉达 - 实体类:

@Entity
@Table(name = "Prada")
@XmlRootElement
@Indexed
@AnalyzerDef(name="customanalyzer", tokenizer = @TokenizerDef(factory = StandardTokenizerFactory.class), 
    filters = { 
        @TokenFilterDef(factory=ISOLatin1AccentFilterFactory.class),
        @TokenFilterDef(factory=LowerCaseFilterFactory.class)})
public class Prada implements Serializable {
 private static final long serialVersionUID = 1L;
@Id
@Basic(optional = false)
@Column(name = "ID")
private Long id;

@Fields({ @Field(index = Index.YES, store = Store.NO), @Field(name = "PradaName_for_sort", index = Index.YES, analyzer = @Analyzer(definition = "customanalyzer")) })
@Column(name = "NAME", length = 100)
private String name;

public Prada () {
}

public Prada (Long id) {
    this.id = id;
}

public Prada (Long id) {
    this.id = id;

}

public Long getId() {
    return id;
}

public void setId(Long id) {
    this.id = id;
}



public String getName() {
    return name;
}

public void setName(String name) {
    this.name = name;
}


@Override
public String toString() {
    return "com.Prac.Prada[ id=" + id + " ]";
}

}

在某个地方找到了这个analyzerDef解决方案但是没有为我工作。有人能为我提供解决方案吗?

主要代码:

  FullTextEntityManager ftem = Search.getFullTextEntityManager(factory.createEntityManager());
  QueryBuilder qb = ftem.getSearchFactory().buildQueryBuilder().forEntity( Prada.class ).get();
  org.apache.lucene.search.Query query = qb.all().getQuery(); 
  FullTextQuery fullTextQuery = ftem.createFullTextQuery(query, Prada.class);
  fullTextQuery.setSort(new Sort(new SortField("PradaName_for_sort", SortField.STRING, true)));
  fullTextQuery.setFirstResult(0).setMaxResults(150);
  int size = fullTextQuery.getResultSize();
  List<Prada> result = fullTextQuery.getResultList();
  for (Pradauser : result) {
    logger.info("Prada Name:" + user.getName());
  }

以下是Lucene的版本(我无法更改):

 <hibernate.version>4.2.8.Final</hibernate.version>
    <hibernate.search.version>4.3.0.Final</hibernate.search.version>

  <dependency>
        <groupId>org.hibernate</groupId>
        <artifactId>hibernate-entitymanager</artifactId>
        <version>4.2.8.Final</version>
    </dependency>
<dependency>
        <groupId>org.apache.lucene</groupId>
        <artifactId>lucene-core</artifactId>
        <version>3.6.2</version>
    </dependency>
    <dependency>
        <groupId>org.apache.lucene</groupId>
        <artifactId>lucene-analyzers</artifactId>
        <version>3.6.2</version>
    </dependency>

更新代码:

@AnalyzerDef(name = "customanalyzer",
tokenizer = @TokenizerDef(factory = KeywordTokenizerFactory.class),
filters = {
    @TokenFilterDef(factory = ASCIIFoldingFilterFactory.class),
    @TokenFilterDef(factory = LowerCaseFilterFactory.class),
    @TokenFilterDef(factory = PatternReplaceFilterFactory.class, params = {
        @Parameter(name = "pattern", value = "('-&\\.,\\(\\))"),
        @Parameter(name = "replacement", value = " "),
        @Parameter(name = "replace", value = "all")
    }),
    @TokenFilterDef(factory = PatternReplaceFilterFactory.class, params = {
        @Parameter(name = "pattern", value = "([^0-9\\p{L} ])"),
        @Parameter(name = "replacement", value = ""),
        @Parameter(name = "replace", value = "all")
    }),
    @TokenFilterDef(factory = TrimFilterFactory.class)
}
)
public class Prada implements Serializable {

@Fields({ @Field(index = Index.YES, store = Store.YES), @Field(name = "PradaName_for_sort", index = Index.YES, analyzer = @Analyzer(definition = "customanalyzer")) })
@Column(name = "NAME", length = 100)
private String name;

1 个答案:

答案 0 :(得分:1)

永远不要使用执行标记化的标记生成器进行排序。您需要使用KeywordTokenizer来确保令牌保持原样。

这是我们用于在我以前的公司进行分类的分析器:

    @AnalyzerDef(name = "TEXT_SORT",
        tokenizer = @TokenizerDef(factory = KeywordTokenizerFactory.class),
        filters = {
                @TokenFilterDef(factory = ASCIIFoldingFilterFactory.class),
                @TokenFilterDef(factory = LowerCaseFilterFactory.class),
                @TokenFilterDef(factory = PatternReplaceFilterFactory.class, params = {
                    @Parameter(name = "pattern", value = "('-&\\.,\\(\\))"),
                    @Parameter(name = "replacement", value = " "),
                    @Parameter(name = "replace", value = "all")
                }),
                @TokenFilterDef(factory = PatternReplaceFilterFactory.class, params = {
                    @Parameter(name = "pattern", value = "([^0-9\\p{L} ])"),
                    @Parameter(name = "replacement", value = ""),
                    @Parameter(name = "replace", value = "all")
                }),
                @TokenFilterDef(factory = TrimFilterFactory.class)
        }
    )

这是最新版本的Hibernate Search,因此您需要对其进行调整。显然,你需要一个s / ASCIIFoldingFilterFactory / ISOLatin1AccentFilterFactory /但是我不确定在3.6.2中是否已经存在PatternReplaceFilterFactory。