Separating Unicode ligature characters

Tags:

Throughout the vast number of unicode characters, there are some that actually represent more than one character, like the U+FB00 ligature ﬀ for two 'f' characters. Is there any way easy to convert characters like these into multiple single characters? Preferably something available in the standard Java API, but I can refer to an external library if need be.

334

asked Aug 24 '11 06:08

nonoitall

2 Answers

U+FB00 is a compatibility character. Normally Unicode doesn't support separate codepoints for ligatures (arguing that it's a layout decision if and when a ligature should be used and should not influence how the data is stored). A few of those still exist to allow round-trip conversion compatibility with older encodings that do represent ligatures as separate entities.

Luckily, the information which characters the ligature represents is present in the Unicode data file and most capable string handling systems have that data built-in.

In Java, you'll need to use the Normalizer class and the NFKC form:

String ff ="\uFB00";
String normalized = Normalizer.normalize(ff, Form.NFKC);
System.out.println(ff + " = " + normalized);

This will print

ﬀ = ff

143

answered Sep 22 '22 16:09

Joachim Sauer

The process you are talking about is called Normalization and is specified in the Unicode Normalization Forms technical note.

There is a class in the Java SE class library called java.text.Normalizer which implements this process. However, you need to read the Unicode document linked above to figure out which of the "normalization forms" you need to use to get the result you want. It is not straightforward ....

answered Sep 20 '22 16:09

Stephen C

Related questions
                            
                                Detect whether an UIImage is PNG or JPEG?
                            
                                Missing Table When Running Django Unittest with Sqlite3
                            
                                How to modify javascript code at run time?
                            
                                Is '\0' guaranteed to be 0?
                            
                                How to create a program to list all the USB devices in a Mac?
                            
                                C# Excel Interop: How to format cells to store values as text
                            
                                Does Content Security Policy block bookmarklets?
                            
                                find size of derived class object using base class pointer
                            
                                how to perform a segue
                            
                                How to make CMake target executed whether specified file was changed?
                            
                                Decoding an input stream
                            
                                How to exclude a directory in a recursive search using grep?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With