有效地替换字符串数组中的字符串

时间:2018-03-07 16:17:42

标签: c# linq

我有一个充满guids的字符串数组。我试图用不同的guid替换某些guid。我的方法如下;

var newArray = this.to.Select(s => s.Replace("e77f75b7-2373-dc11-8f13-0019bb2ca0a0", "1fe8f3f6-fe17-e811-80d8-00155d5ce473")
    .Replace("fbd0c892-2373-dc11-8f13-0019bb2ca0a0", "1fe8f3f6-fe17-e811-80d8-00155d5ce473")
    .Replace("76cd4297-1e31-dc11-95d8-0019bb2ca0a0", "eb892fb0-fe17-e811-80d8-00155d5ce473")
    .Replace("cd42bb68-2073-dc11-8f13-0019bb2ca0a0", "dc6077e2-fe17-e811-80d8-00155d5ce473")
    .Replace("96b97150-cd45-e111-a3d5-00155d10010f", "1fe8f3f6-fe17-e811-80d8-00155d5ce473")
    ).ToArray();

我有一些我正在做的字段,它导致OutOfMemoryException。是因为Replace()方法每次都在创建一个新数组吗?有一个更有效的方法来做一个字符串数组?这个方法正在运行成千上万的记录,所以我认为这是个问题。当我评论这些行时,我没有得到例外。

编辑:'到'中的数据变量在每种情况下都是一个短字符串,但这是为数千条记录运行的。所以'到'对于一条记录可能看起来像这样;

"systemuser|76cd4297-1e31-dc11-95d8-0019bb2ca0a0;contact|96b97150-cd45-e111-a3d5-00155d10010f"

它可能有我要替换的任何guid,所以即使它可能只有一个guid用于该记录,我需要运行完整的替换()以防万一它有任何一个在它。

任何指针都会很棒!感谢。

3 个答案:

答案 0 :(得分:3)

我会使用替换词典 - 它更容易维护,更容易理解(我认为)所以它更容易:

Boilerplate并创建演示数据/替换dict:

using System;
using System.Collections.Generic;
using System.Data;
using System.Linq;

internal class Program
{
    static void Main(string[] args)
    {
        // c#7 inline func
        string[] CreateDemoData(Dictionary<string, string> replDict)
        {
            // c#7 inline func
            string FilText(string s) => $"Some text| that also incudes; {s} and more.";

            return Enumerable
                .Range(1, 5)
                .Select(i => FilText(Guid.NewGuid().ToString()))
                .Concat(replDict.Keys.Select(k => FilText(k)))
                .OrderBy(t => Guid.NewGuid().GetHashCode())
                .ToArray();
        }

        // replacement dict
        var d = new Dictionary<string, string>
        {
            ["e77f75b7-2373-dc11-8f13-0019bb2ca0a0"] = "e77f75b7-replaced",
            ["fbd0c892-2373-dc11-8f13-0019bb2ca0a0"] = "fbd0c892-replaced",
            ["76cd4297-1e31-dc11-95d8-0019bb2ca0a0"] = "76cd4297-replaced",
            ["cd42bb68-2073-dc11-8f13-0019bb2ca0a0"] = "cd42bb68-replaced",
            ["96b97150-cd45-e111-a3d5-00155d10010f"] = "96b97150-replaced",
        };

        var arr = CreateDemoData(d);

创建实际替换数组的代码:

        // c#7 inline func
        string Replace(string a, Dictionary<string, string> dic)
        {
            foreach (var key in dic.Keys.Where(k => a.Contains(k)))
                a = a.Replace(key, dic[key]);

            return a;
        }

        // select value from dict in key in dict else leave unmodified            
        var b = arr.Select(a => Replace(a, d));
        // if you have really that much data (20k guids of ~50byte length
        // is not really much imho) you can use the same approach for in
        // place replacement - just foreach over your array.

输出代码:

        Console.WriteLine("\nBefore:");
        foreach (var s in arr)
            Console.WriteLine(s);

        Console.WriteLine("\nAfter:");
        foreach (var s in b)
            Console.WriteLine(s);

        Console.ReadLine(); 
    }
}

输出:

Before:
Some text| that also incudes; a5ceefd8-1388-47cd-b69e-55b6ddbbc133 and more.
Some text| that also incudes; 76cd4297-1e31-dc11-95d8-0019bb2ca0a0 and more.
Some text| that also incudes; 3311a8c5-015e-4260-af80-86b20b277234 and more.
Some text| that also incudes; ed10c79c-dad6-4c88-865c-4d7624945d66 and more.
Some text| that also incudes; 96b97150-cd45-e111-a3d5-00155d10010f and more.
Some text| that also incudes; 0226d9b1-c5f0-41fb-9294-bc9297e8afd9 and more.
Some text| that also incudes; e77f75b7-2373-dc11-8f13-0019bb2ca0a0 and more.
Some text| that also incudes; a04d1e34-e7bc-4bbc-ae0e-12ec846a353c and more.
Some text| that also incudes; cd42bb68-2073-dc11-8f13-0019bb2ca0a0 and more.
Some text| that also incudes; fbd0c892-2373-dc11-8f13-0019bb2ca0a0 and more.

输出:

After:
Some text| that also incudes; a5ceefd8-1388-47cd-b69e-55b6ddbbc133 and more.
Some text| that also incudes; 76cd4297-replaced and more.
Some text| that also incudes; 3311a8c5-015e-4260-af80-86b20b277234 and more.
Some text| that also incudes; ed10c79c-dad6-4c88-865c-4d7624945d66 and more.
Some text| that also incudes; 96b97150-replaced and more.
Some text| that also incudes; 0226d9b1-c5f0-41fb-9294-bc9297e8afd9 and more.
Some text| that also incudes; e77f75b7-replaced and more.
Some text| that also incudes; a04d1e34-e7bc-4bbc-ae0e-12ec846a353c and more.
Some text| that also incudes; cd42bb68-replaced and more.
Some text| that also incudes; fbd0c892-replaced and more.

答案 1 :(得分:0)

我通过使用正则表达式提取字段,然后使用替换词典来应用更改,然后重新构建字符串,在单次扫描中执行此操作:

IDictionary<string, string> replacements = new Dictionary<string, string>
{
    {"76cd4297-1e31-dc11-95d8-0019bb2ca0a0","something else"},
    //etc
};
var newData = data
    //.AsParallel() //for speed
    .Select(d => Regex.Match(d, @"^(?<f1>[^\|]*)\|(?<f2>[^;]*);(?<f3>[^\|]*)\|(?<f4>.*)$"))
    .Where(m => m.Success)
    .Select(m => new
    {
        field1 = m.Groups["f1"].Value,
        field2 = m.Groups["f2"].Value,
        field3 = m.Groups["f3"].Value,
        field4 = m.Groups["f4"].Value
    })
    .Select(x => new
    {
        x.field1,
        field2 = replacements.TryGetValue(x.field2, out string r2) ? r2 : x.field2,
        x.field3,
        field4 = replacements.TryGetValue(x.field4, out string r4) ? r4 : x.field4
    })
    .Select(x => $"{x.field1}|{x.field2};{x.field3}|{x.field4}")
    .ToArray();

答案 2 :(得分:-1)

您是否使用StringBuilder进行了测试?

StringBuilder sb = new StringBuilder(string.Join(",", this.to));

      string tempStr = sb
            .Replace("e77f75b7-2373-dc11-8f13-0019bb2ca0a0", "1fe8f3f6-fe17-e811-80d8-00155d5ce473")
            .Replace("fbd0c892-2373-dc11-8f13-0019bb2ca0a0", "1fe8f3f6-fe17-e811-80d8-00155d5ce473")
            .Replace("76cd4297-1e31-dc11-95d8-0019bb2ca0a0", "eb892fb0-fe17-e811-80d8-00155d5ce473")
            .Replace("cd42bb68-2073-dc11-8f13-0019bb2ca0a0", "dc6077e2-fe17-e811-80d8-00155d5ce473")
            .Replace("96b97150-cd45-e111-a3d5-00155d10010f", "1fe8f3f6-fe17-e811-80d8-00155d5ce473")
            .ToString();

      var newArray = tempStr.Split(',');