I have a bunch of names, and I want to obtain the unique names. However, due to spelling errors and inconsistencies in the data the names might be written down wrong. I am looking for a way to check in a vector of strings if two of them are similair. For example: <pre class="prettyprint"><code>pres <- c(" Obama, B.","Bush, G.W.","Obama, B.H.","Clinton, W.J.") </code></pre> I want to find that <code>" Obama, B."</code> and <code>"Obama, B.H."</code> are very similar. Is there a way to do this?

This can be done based on eg the Levenshtein distance. There are multiple implementations of this in different packages. Some solutions and packages can be found in the answers of these questions: <ul> <li>agrep: only return best match(es)</li> <li>In R, how do I replace a string that contains a certain pattern with another string?</li> <li>Fast Levenshtein distance in R?</li> </ul> But most often <code>agrep</code> will do what you want : <pre class="prettyprint"><code>> sapply(pres,agrep,pres) $` Obama, B.` [1] 1 3 $`Bush, G.W.` [1] 2 $`Obama, B.H.` [1] 1 3 $`Clinton, W.J.` [1] 4 </code></pre>

How to measure similarity between strings?

Tags:

string

regex

r

r-faq

I have a bunch of names, and I want to obtain the unique names. However, due to spelling errors and inconsistencies in the data the names might be written down wrong. I am looking for a way to check in a vector of strings if two of them are similair.

For example:

pres <- c(" Obama, B.","Bush, G.W.","Obama, B.H.","Clinton, W.J.")

I want to find that " Obama, B." and "Obama, B.H." are very similar. Is there a way to do this?

805

asked May 18 '11 11:05

Sacha Epskamp

1 Answers

This can be done based on eg the Levenshtein distance. There are multiple implementations of this in different packages. Some solutions and packages can be found in the answers of these questions:

agrep: only return best match(es)
In R, how do I replace a string that contains a certain pattern with another string?
Fast Levenshtein distance in R?

But most often agrep will do what you want :

> sapply(pres,agrep,pres) $` Obama, B.` [1] 1 3  $`Bush, G.W.` [1] 2  $`Obama, B.H.` [1] 1 3  $`Clinton, W.J.` [1] 4

176

answered Oct 03 '22 10:10

Joris Meys

Related questions
                            
                                Reuse part of a Regex pattern
                            
                                Is there a good way to replace home directory with tilde in bash?
                            
                                Why doesn't a non-greedy quantifier sometimes work in Oracle regex?
                            
                                Regex - match everything but forward slash
                            
                                Regex with -, ::, ( and )
                            
                                Escaping a parenthesis in grep/ack
                            
                                How can you detect if two regular expressions overlap in the strings they can match?
                            
                                How to make "grep" read patterns from a file?
                            
                                Convert SRE_Match object to string
                            
                                How do you capture a group with regex?
                            
                                Pattern matching field names with jq
                            
                                What does (?i) and ?@ in this regex mean [duplicate]
                            
                                Regex BNF Grammar
                            
                                php preg_match non greedy? [duplicate]
                            
                                Regex in spring controller
                            
                                Regular Expression Wildcard Matching
                            
                                Test if a string is regex
                            
                                RegEx for valid international mobile phone number [duplicate]
                            
                                Regex - check if input still has chances to become matching
                            
                                Python regex: matching a parenthesis within parenthesis

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With