Standard <code>grep</code>/<code>pcregrep</code> etc. can conveniently be used with binary files for ASCII or UTF8 data - is there a simple way to make them try UTF16 too (preferably simultaneously, but instead will do)? Data I'm trying to get is all ASCII anyway (references in libraries etc.), it just doesn't get found as sometimes there's 00 between any two characters, and sometimes there isn't. I don't see any way to get it done semantically, but these 00s should do the trick, except I cannot easily use them on command line.

The easiest way is to just convert the text file to utf-8 and pipe that to grep: <pre class="prettyprint"><code>iconv -f utf-16 -t utf-8 file.txt | grep query </code></pre> I tried to do the opposite (convert my query to utf-16) but it seems as though grep doesn't like that. I think it might have to do with endianness, but I'm not sure. It seems as though grep will convert a query that is utf-16 to utf-8/ascii. Here is what I tried: <pre class="prettyprint"><code>grep `echo -n query | iconv -f utf-8 -t utf-16 | sed 's/..//'` test.txt </code></pre> If test.txt is a utf-16 file this won't work, but it does work if test.txt is ascii. I can only conclude that grep is converting my query to ascii. EDIT: Here's a really really crazy one that kind of works but doesn't give you very much useful info: <pre class="prettyprint"><code>hexdump -e '/1 "%02x"' test.txt | grep -P `echo -n Test | iconv -f utf-8 -t utf-16 | sed 's/..//' | hexdump -e '/1 "%02x"'` </code></pre> How does it work? Well it converts your file to hex (without any extra formatting that hexdump usually applies). It pipes that into grep. Grep is using a query that is constructed by echoing your query (without a newline) into iconv which converts it to utf-16. This is then piped into sed to remove the BOM (the first two bytes of a utf-16 file used to determine endianness). This is then piped into hexdump so that the query and the input are the same. Unfortunately I think this will end up printing out the ENTIRE file if there is a single match. Also this won't work if the utf-16 in your binary file is stored in a different endianness than your machine. EDIT2: Got it!!!! <pre class="prettyprint"><code>grep -P `echo -n "Test" | iconv -f utf-8 -t utf-16 | sed 's/..//' | hexdump -e '/1 "x%02x"' | sed 's/x/\\\\x/g'` test.txt </code></pre> This searches for the hex version of the string <code>Test</code> (in utf-16) in the file <code>test.txt</code>

I found the below solution worked best for me, from https://www.splitbits.com/2015/11/11/tip-grep-and-unicode/ Grep does not play well with Unicode, but it can be worked around. For example, to find, <pre class="prettyprint"><code>Some Search Term </code></pre> in a UTF-16 file, use a regular expression to ignore the first byte in each character, <pre class="prettyprint"><code>S.o.m.e. .S.e.a.r.c.h. .T.e.r.m </code></pre> Also, tell grep to treat the file as text, using '-a', the final command looks like this, <pre class="prettyprint"><code>grep -a 'S.o.m.e. .S.e.a.r.c.h. .T.e.r.m' utf-16-file.txt </code></pre>

grepping binary files and UTF16

2 Answers

The easiest way is to just convert the text file to utf-8 and pipe that to grep:

iconv -f utf-16 -t utf-8 file.txt | grep query

I tried to do the opposite (convert my query to utf-16) but it seems as though grep doesn't like that. I think it might have to do with endianness, but I'm not sure.

It seems as though grep will convert a query that is utf-16 to utf-8/ascii. Here is what I tried:

grep `echo -n query | iconv -f utf-8 -t utf-16 | sed 's/..//'` test.txt

If test.txt is a utf-16 file this won't work, but it does work if test.txt is ascii. I can only conclude that grep is converting my query to ascii.

EDIT: Here's a really really crazy one that kind of works but doesn't give you very much useful info:

hexdump -e '/1 "%02x"' test.txt | grep -P `echo -n Test | iconv -f utf-8 -t utf-16 | sed 's/..//' | hexdump -e '/1 "%02x"'`

How does it work? Well it converts your file to hex (without any extra formatting that hexdump usually applies). It pipes that into grep. Grep is using a query that is constructed by echoing your query (without a newline) into iconv which converts it to utf-16. This is then piped into sed to remove the BOM (the first two bytes of a utf-16 file used to determine endianness). This is then piped into hexdump so that the query and the input are the same.

Unfortunately I think this will end up printing out the ENTIRE file if there is a single match. Also this won't work if the utf-16 in your binary file is stored in a different endianness than your machine.

EDIT2: Got it!!!!

grep -P `echo -n "Test" | iconv -f utf-8 -t utf-16 | sed 's/..//' | hexdump -e '/1 "x%02x"' | sed 's/x/\\\\x/g'` test.txt

This searches for the hex version of the string Test (in utf-16) in the file test.txt

answered Sep 30 '22 09:09

Niki Yoshiuchi

I found the below solution worked best for me, from https://www.splitbits.com/2015/11/11/tip-grep-and-unicode/

Grep does not play well with Unicode, but it can be worked around. For example, to find,

Some Search Term

in a UTF-16 file, use a regular expression to ignore the first byte in each character,

S.o.m.e. .S.e.a.r.c.h. .T.e.r.m

Also, tell grep to treat the file as text, using '-a', the final command looks like this,

grep -a 'S.o.m.e. .S.e.a.r.c.h. .T.e.r.m' utf-16-file.txt

answered Sep 30 '22 10:09

nirmal

Related questions
                            
                                Why is this LSEP symbol showing up on Chrome and not Firefox or Edge?
                            
                                Python regex matching Unicode properties
                            
                                How does uʍop-ǝpᴉsdn text work?
                            
                                How to match Cyrillic characters with a regular expression
                            
                                Regular Expression Arabic characters and numbers only
                            
                                How to get rid of non-ascii characters in ruby
                            
                                removing emojis from a string in Python
                            
                                Regex to match Egyptian Hieroglyphics [closed]
                            
                                Should I use accented characters in URLs?
                            
                                Can UTF-8 contain zero byte?
                            
                                Are email addresses allowed to contain non-alphanumeric characters?
                            
                                Difference between MBCS and UTF-8 on Windows
                            
                                Font Awesome & Unicode
                            
                                SQLite, python, unicode, and non-utf data
                            
                                How do I remove the BOM character from my xml file [duplicate]
                            
                                Python NLTK: SyntaxError: Non-ASCII character '\xc3' in file (Sentiment Analysis -NLP)
                            
                                Fixing broken UTF-8 encoding
                            
                                How can I get a Unicode character's code?
                            
                                Why does the size of this Python String change on a failed int conversion
                            
                                How do I turn off Unicode in a VC++ project?

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

grepping binary files and UTF16

Tags:

grep

unicode

utf-16

taw

People also ask

2 Answers

Niki Yoshiuchi

nirmal

Recent Activity

Donate For Us