Logo Questions Linux Laravel Mysql Ubuntu Git Menu
 

Generic way to simplify string, remove diacritics

I'm searching for a method to remove diacritics and other letter marks in a text and simplify it in a way that it is a good fit for a text search index.

For removing the diacritics, I already found these:

  • questions for PHP: 1, 2
  • question for Java: 1, related: 2
  • question for Bash: 1
  • questions for .Net: 1, 2
  • question for Javascript: 1
  • question for Python: 1

I was wondering about a generic solution, language independent. (Also, this reference list might be useful for some.)

Removing the diacritics works for äöüò, etc. But I also want:

  • ø → o
  • Я → R
  • Ł → L
  • ɲ → n
  • æ → a (it could also be "ae" but in my case, "a" makes more sense because I also want to replace "ae" by "a")

For example, I want to index the name Røyksopp which sometimes also occurs as Röyksopp just under the simplified name Royksopp. Or KoЯn should be KoRn.

like image 620
Albert Avatar asked Aug 25 '26 15:08

Albert


1 Answers

Some ICU magic:

echo "ë ö ø Я Ł ɲ æ å ñ 開 당" | uconv -x any-name | perl -wpne 's/ WITH [^}]+//g;' | uconv -x name-any | uconv -x any-latin -t iso-8859-1 -c | uconv -f iso-8859-1 -t ascii -x latin-ascii -c

yields

e o o A L n ae a n ki dang

This uses the cmdline tool uconv, but the same can be done with ICU's Java or C or C++ API, and ICU has bindings for almost any language.

Note Я -> A because that is the correct behavior. What you want is not how Unicode defines that character - blame KoЯn for abusing it.

like image 173
Tino Didriksen Avatar answered Aug 28 '26 05:08

Tino Didriksen



Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!