I have some text that contains Unicode escape sequences like \u003C. This is what I came up with to unescape it: <code>string.gsub(/\u(....)/) {|m| [$1].pack("H*").unpack("n*").pack("U*")}</code> Is it correct? (i.e. it seems to work with my tests, but can someone more knowledgeable find a problem with it?)

Your regex, <code>/\u(....)/</code>, has some problems. First of all, <code>\u</code> doesn't work the way you think it does, in 1.9 you'll get an error and in 1.8 it will just match a single <code>u</code> rather than the <code>\u</code> pair that you're looking for; you should use <code>/\\u/</code> to find the literal <code>\u</code> that you want. Secondly, your <code>(....)</code> group is much too permissive, that will allow any four characters through and that's not what you want. In 1.9, you want <code>(\h{4})</code> (four hexadecimal digits) but in 1.8 you'd need <code>([\da-fA-F]{4})</code> as <code>\h</code> is a new thing. So if you want your regex to work in both 1.8 and 1.9, you should use <code>/\\u([\da-fA-F]{4})/</code>. This gives you the following in 1.8 and 1.9: <pre class="prettyprint"><code>>> s = 'Where is \u03bc pancakes \u03BD house? And u1123!' => "Where is \\u03bc pancakes \\u03BD house? And u1123!" >> s.gsub(/\\u([\da-fA-F]{4})/) {|m| [$1].pack("H*").unpack("n*").pack("U*")} => "Where is μ pancakes ν house? And u1123!" </code></pre> Using <code>pack</code> and <code>unpack</code> to mangle the hex number into a Unicode character is probably good enough but there may be better ways.

Is this the best way to unescape unicode escape sequences in Ruby?

1 Answers

Your regex, /\u(....)/, has some problems.

First of all, \u doesn't work the way you think it does, in 1.9 you'll get an error and in 1.8 it will just match a single u rather than the \u pair that you're looking for; you should use /\\u/ to find the literal \u that you want.

Secondly, your (....) group is much too permissive, that will allow any four characters through and that's not what you want. In 1.9, you want (\h{4}) (four hexadecimal digits) but in 1.8 you'd need ([\da-fA-F]{4}) as \h is a new thing.

So if you want your regex to work in both 1.8 and 1.9, you should use /\\u([\da-fA-F]{4})/. This gives you the following in 1.8 and 1.9:

>> s = 'Where is \u03bc pancakes \u03BD house? And u1123!'
=> "Where is \\u03bc pancakes \\u03BD house? And u1123!"
>> s.gsub(/\\u([\da-fA-F]{4})/) {|m| [$1].pack("H*").unpack("n*").pack("U*")}
=> "Where is μ pancakes ν house? And u1123!"

Using pack and unpack to mangle the hex number into a Unicode character is probably good enough but there may be better ways.

answered Dec 02 '22 21:12

mu is too short

Related questions
                            
                                Bytes vs codepoints in ruby
                            
                                Spring stopping Rails console from running
                            
                                How can I use C# style enumerations in Ruby?
                            
                                How does Ruby know where to find a required file?
                            
                                how to parse multivalued field from URL query in Rails
                            
                                "k.send :hello" - if k is the "receiver", who is the sender?
                            
                                How to combine ActiveRecord objects?
                            
                                Ruby installer on Windows 7 64-bit machine
                            
                                Ruby: Is there something like Enumerable#drop that returns an enumerator instead of an array?
                            
                                Ordering .each results in the view
                            
                                Are there any solutions for translating measurement units on Rails?
                            
                                Cucumber test order of elements in table
                            
                                get div nested in div element using Nokogiri
                            
                                Difference between p in a rails view and puts
                            
                                call a specific url with rspec
                            
                                Trouble yielding inside a block/lambda
                            
                                Haml "Illegal Nesting" problem; how to place multiple code elements in the same tag?
                            
                                Create a trial period for ruby on rails web app
                            
                                ruby/rails equivalent to javascript decodeURIComponent?
                            
                                Stripping commas from Integers or decimals in rails

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

Is this the best way to unescape unicode escape sequences in Ruby?

Tags:

ruby

unicode

Eric Mason

People also ask

1 Answers

mu is too short

Recent Activity

Donate For Us