I have a list of strings (company names, in this case), and a Java program that extracts a list of things that look like company names out of mostly-unstructured text. I need to match each element of extracted text to a string in the list. Caveat: the unstructured text has typos, things like "Blah, Inc." referred to as "Blah," etc. I've tried Levenshtein Edit Distance, but that fails for predictable reasons. Are there known best-practices ways of tackling this problem? Or am I back to manual data-entry?
You might want to have a look at Apache Stanbol, it plugs together NER engines (I think one is based on a gazetteer you supply) and linking engines to resolve your detected entities. I haven't used it myself and it's still in incubation, but might suit what you're looking for.
There is also a bit of research in this space in the TAC Knowledge Base Population track (Entity Linking). The task pops-up in different places and you should also have luck in conferences like ACL, EMNLP, SIGIR, etc (this list is by no means complete).
The TAC systems link to a subset of Wikipedia, which might help with your name variation since pages have "redirects", which are essentially aliases for a particular page.
For example, following pages redirect to "Apple Inc.", but you probably want to extract the redirects either from a raw Wikipedia dump or from a clean source like DBPedia or Freebase.
This is not a simple problem, and there are entire companies built around trying to solve it (even for reduced matching sets like company names versus the general case).
If you can identify a discrete number of patterns that valid company names fall into, and that noise does not fall into, then you could tackle this with a series of regular expression matches.
If the patterns are difficult or too numerous, then you could try developing a probabilistic model, perhaps something like a Bayesian network. You would take a subset of your data for training, and perhaps a second subset for a quick validation, and grow the network. Techniques might include genetic programming or setting up a neural network. This approach is obviously not lightweight, and you'd probably want to consider your need carefully before going down this road.
If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!
Donate Us With