Lucene.Net(4.8)自动完成/自动建议

时间:2020-02-26 15:53:23

标签: c# lucene.net

我想使用Lucene.Net 4.8实现可搜索的索引,该索引可以为用户提供针对单个单词和短语的建议/自动完成功能。

索引已成功创建;这些建议是我拖延的地方。

版本4.8似乎引入了许多重大更改,而我发现没有可用的示例可用。

我站在哪里

LuceneVersion是这样的:

private readonly LuceneVersion LuceneVersion = LuceneVersion.LUCENE_48;

解决方案1 ​​

I've tried this,但不能越过reader.Terms

    public void TryAutoComplete()
    {
        var analyzer = new EnglishAnalyzer(LuceneVersion);
        var config = new IndexWriterConfig(LuceneVersion, analyzer);
        RAMDirectory dir = new RAMDirectory();
        using (IndexWriter iw = new IndexWriter(dir, config))
        {
            Document d = new Document();
            TextField f = new TextField("text","",Field.Store.YES);
            d.Add(f);
            f.SetStringValue("abc");
            iw.AddDocument(d);
            f.SetStringValue("colorado");
            iw.AddDocument(d);
            f.SetStringValue("coloring book");
            iw.AddDocument(d);
            iw.Commit();
            using (IndexReader reader = iw.GetReader(false))
            {
                TermEnum terms = reader.Terms(new Term("text", "co"));
                int maxSuggestsCpt = 0;
                // will print:
                // colorado
                // coloring book
                do
                {
                    Console.WriteLine(terms.Term.Text);
                    maxSuggestsCpt++;
                    if (maxSuggestsCpt >= 5)
                        break;
                }
                while (terms.Next() && terms.Term.Text.StartsWith("co"));
            }
        }
    }

reader.Terms no longer exists。作为Lucene的新手,目前尚不清楚如何重构它。

解决方案2

尝试this时,我抛出了错误:

    public void TryAutoComplete2()
    {
        using(var analyzer = new EnglishAnalyzer(LuceneVersion))
        {
            IndexWriterConfig config = new IndexWriterConfig(LuceneVersion, analyzer);
            RAMDirectory dir = new RAMDirectory();
            using(var iw = new IndexWriter(dir,config))
            {
                Document d = new Document()
                {
                    new TextField("text", "this is a document with a some words",Field.Store.YES),
                    new Int32Field("id", 42, Field.Store.YES)
                };

                iw.AddDocument(d);
                iw.Commit();

                using (IndexReader reader = iw.GetReader(false))
                using (SpellChecker speller = new SpellChecker(new RAMDirectory()))
                {
                    //ERROR HERE!!!
                    speller.IndexDictionary(new LuceneDictionary(reader, "text"), config, false);
                    string[] suggestions = speller.SuggestSimilar("dcument", 5);
                    IndexSearcher searcher = new IndexSearcher(reader);
                    foreach (string suggestion in suggestions)
                    {
                        TopDocs docs = searcher.Search(new TermQuery(new Term("text", suggestion)), null, Int32.MaxValue);
                        foreach (var doc in docs.ScoreDocs)
                        {
                            System.Diagnostics.Debug.WriteLine(searcher.Doc(doc.Doc).Get("id"));
                        }
                    }
                }
            }
        }
    }

调试时,speller.IndexDictionary(new LuceneDictionary(reader, "text"), config, false);会引发The object cannot be set twice!错误,我无法解释。

欢迎任何想法。

澄清

我想返回给定输入的建议术语列表,而不是文档或其全部内容。

例如,如果文档包含“你好,我的名字叫克拉克。我来自亚特兰大”,而我提交了“ Atl”,那么“ Atlanta”应该作为建议。

1 个答案:

答案 0 :(得分:1)

如果我对您的理解正确,则可能会使索引设计有些复杂。如果您的目标是使用Lucene自动完成完成,则希望创建一个您认为完成的术语的索引。然后只需使用带有部分单词或短语的PrefixQuery来查询索引。

using Lucene.Net.Analysis;
using Lucene.Net.Analysis.En;
using Lucene.Net.Documents;
using Lucene.Net.Index;
using Lucene.Net.Search;
using Lucene.Net.Store;
using Lucene.Net.Util;
using System;
using System.Linq;

namespace LuceneDemoApp
{
    class LuceneAutoCompleteIndex : IDisposable
    {
        const LuceneVersion Version = LuceneVersion.LUCENE_48;
        RAMDirectory Directory;
        Analyzer Analyzer;
        IndexWriterConfig WriterConfig;

        private void IndexDoc(IndexWriter writer, string term)
        {
            Document doc = new Document();
            doc.Add(new StringField(FieldName, term, Field.Store.YES));
            writer.AddDocument(doc);
        }

        public LuceneAutoCompleteIndex(string fieldName, int maxResults)
        {
            FieldName = fieldName;
            MaxResults = maxResults;
            Directory = new RAMDirectory();
            Analyzer = new EnglishAnalyzer(Version);
            WriterConfig = new IndexWriterConfig(Version, Analyzer);
            WriterConfig.OpenMode = OpenMode.CREATE_OR_APPEND;
        }

        public string FieldName { get; }
        public int MaxResults { get; set; }

        public void Add(string term)
        {
            using (var writer = new IndexWriter(Directory, WriterConfig))
            {
                IndexDoc(writer, term);
            }
        }

        public void AddRange(string[] terms)
        {
            using (var writer = new IndexWriter(Directory, WriterConfig))
            {
                foreach (string term in terms)
                {
                    IndexDoc(writer, term);
                }
            }
        }

        public string[] WhereStartsWith(string term)
        {
            using (var reader = DirectoryReader.Open(Directory))
            {
                IndexSearcher searcher = new IndexSearcher(reader);
                var query = new PrefixQuery(new Term(FieldName, term));
                TopDocs foundDocs = searcher.Search(query, MaxResults);
                var matches = foundDocs.ScoreDocs
                    .Select(scoreDoc => searcher.Doc(scoreDoc.Doc).Get(FieldName))
                    .ToArray();

                return matches;
            }
        }

        public void Dispose()
        {
            Directory.Dispose();
            Analyzer.Dispose();
        }
    }
}

运行此:

var indexValues = new string[] { "apple fruit", "appricot", "ape", "avacado", "banana", "pear" };
var index = new LuceneAutoCompleteIndex("fn", 10);
index.AddRange(indexValues);

var matches = index.WhereStartsWith("app");
foreach (var match in matches)
{
    Console.WriteLine(match);
}

您得到了:

apple fruit
appricot