如何在 Python 中对 Levenshtein 距离超过 80% 的词进行分组

Question

假设我有一个列表：-

person_name = ['zakesh', 'oldman LLC', 'bikash', 'goldman LLC', 'zikash','rakesh']

我正在尝试以这种方式对列表进行分组，以便 Levenshtein distance between two strings is maximum. For finding out the ratio between two words, I am using a python package fuzzywuzzy。

例子：-

>>> from fuzzywuzzy import fuzz
>>> combined_list = ['rakesh', 'zakesh', 'bikash', 'zikash', 'goldman LLC', 'oldman LLC']
>>> fuzz.ratio('goldman LLC', 'oldman LLC')
95
>>> fuzz.ratio('rakesh', 'zakesh')
83
>>> fuzz.ratio('bikash', 'zikash')
83
>>>

我的最终目标：

My end goal is to group the words such that Levenshtein distance between them is more than 80 percent?

我的列表应该是这样的：-

person_name = ['bikash', 'zikash', 'rakesh', 'zakesh', 'goldman LLC', 'oldman LLC'] because the distance between `bikash` and `zikash` is very high so they should be together.

代码：

我正在尝试通过排序来实现这一点，但关键函数应该是 fuzz.ratio。好吧，下面的代码不起作用，但我正在从这个角度解决问题。

from fuzzywuzzy import fuzz
combined_list = ['rakesh', 'zakesh', 'bikash', 'zikash', 'goldman LLC', 'oldman LLC']
combined_list.sort(key=lambda x, y: fuzz.ratio(x, y))
print combined_list

Could anyone help me to combine the words so that Levenshtein distance between them is more than 80 percent?

Answer 1

这对名称进行分组

from fuzzywuzzy import fuzz

combined_list = ['rakesh', 'zakesh', 'bikash', 'zikash', 'goldman LLC', 'oldman LLC']
combined_list.append('bakesh')
print('input names:', combined_list)

grs = list() # groups of names with distance > 80
for name in combined_list:
    for g in grs:
        if all(fuzz.ratio(name, w) > 80 for w in g):
            g.append(name)
            break
    else:
        grs.append([name, ])

print('output groups:', grs)
outlist = [el for g in grs for el in g]
print('output list:', outlist)

生产

input names: ['rakesh', 'zakesh', 'bikash', 'zikash', 'goldman LLC', 'oldman LLC', 'bakesh']
output groups: [['rakesh', 'zakesh', 'bakesh'], ['bikash', 'zikash'], ['goldman LLC', 'oldman LLC']]
output list: ['rakesh', 'zakesh', 'bakesh', 'bikash', 'zikash', 'goldman LLC', 'oldman LLC']

如您所见，名称已正确分组，但顺序可能不是您想要的顺序。

如何在 Python 中对 Levenshtein 距离超过 80% 的词进行分组

How to group words whose Levenshtein distance is more than 80 percent in Python

python

fuzzy-search

group-by

fuzzy-logic

levenshtein-distance