How to check whether a file is valid UTF-8?

Tags:

I'm processing some data files that are supposed to be valid UTF-8 but aren't, which causes the parser (not under my control) to fail. I'd like to add a stage of pre-validating the data for UTF-8 well-formedness, but I've not yet found a utility to help do this.

There's a web service at W3C which appears to be dead, and I've found a Windows-only validation tool that reports invalid UTF-8 files but doesn't report which lines/characters to fix.

I'd be happy with either a tool I can drop in and use (ideally cross-platform), or a ruby/perl script I can make part of my data loading process.

801

asked Sep 22 '08 14:09

Ian Dickinson

2 Answers

You can use GNU iconv:

$ iconv -f UTF-8 your_file -o /dev/null; echo $?

Or with older versions of iconv, such as on macOS:

$ iconv -f UTF-8 your_file > /dev/null; echo $?

The command will return 0 if the file could be converted successfully, and 1 if not. Additionally, it will print out the byte offset where the invalid byte sequence occurred.

Edit: The output encoding doesn't have to be specified, it will be assumed to be UTF-8.

182

answered Oct 19 '22 18:10

Torsten Marek

Use python and str.encode|decode functions.

>>> a="γεια" >>> a '\xce\xb3\xce\xb5\xce\xb9\xce\xb1' >>> b='\xce\xb3\xce\xb5\xce\xb9\xff\xb1' # note second-to-last char changed >>> print b.decode("utf_8") Traceback (most recent call last):   File "<stdin>", line 1, in <module>   File "/usr/local/lib/python2.5/encodings/utf_8.py", line 16, in decode     return codecs.utf_8_decode(input, errors, True) UnicodeDecodeError: 'utf8' codec can't decode byte 0xff in position 6: unexpected code byte

The exception thrown has the info requested in its .args property.

>>> try: print b.decode("utf_8") ... except UnicodeDecodeError, exc: pass ... >>> exc UnicodeDecodeError('utf8', '\xce\xb3\xce\xb5\xce\xb9\xff\xb1', 6, 7, 'unexpected code byte') >>> exc.args ('utf8', '\xce\xb3\xce\xb5\xce\xb9\xff\xb1', 6, 7, 'unexpected code byte')

answered Oct 19 '22 18:10

tzot

Related questions
                            
                                Ruby Email validation with regex
                            
                                MVC 3 jQuery Validation/globalizing of number/decimal field
                            
                                How to check if a string is a valid regex in Python?
                            
                                Input type number "only numeric value" validation
                            
                                Method must have signature "String method() ...[etc]..." but has signature "void method()"
                            
                                calling custom validation methods in Rails
                            
                                Date validation with ASP.NET validator
                            
                                Laravel Password & Password_Confirmation Validation
                            
                                How to properly validate input values with React.JS?
                            
                                How to add a RequiredFieldValidator to DropDownList control?
                            
                                Validate the number of has_many items in Ruby on Rails
                            
                                HTML5 'required' validation in Ruby on Rails forms
                            
                                How to add a Not Equal To rule in jQuery.validation
                            
                                Laravel password validation rule
                            
                                DataAnnotations: Recursively validating an entire object graph
                            
                                Email Address Validation for ASP.NET
                            
                                Rails before_validation strip whitespace best practices
                            
                                Spring - Redirect after POST (even with validation errors)
                            
                                How to validate numeric values which may contain dots or commas?
                            
                                Validation of a list of objects in Spring

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With

How to check whether a file is valid UTF-8?

Tags:

validation

utf-8

internationalization

Ian Dickinson

People also ask

2 Answers

Torsten Marek

tzot

Recent Activity

Donate For Us