What programmers need to know about encodings and charsets (2011) (opens in new tab)

(kunststube.net)

70 pointsneiesc5y ago22 comments

22 comments

I was looking for the catch. Here it is: "It's really simple: Know what encoding a certain piece of text, that is, a certain byte sequence, is in, then interpret it with that encoding."

That's like "knowing" the truth. How?

I have received some very interesting files that made Python yack unicode errors, again and again. Why? Not only did I not "know" what encoding it was in -- the encodings changed at different points in the stream of bytes. I call this "slamming bytes together" because somewhere along the line, someone's program did exactly that.

Everything is simple -- until it isn't.

banthar5y ago

There is nothing you can do with text file with unknown encoding but treat it as an array of bytes.

If you start guessing the encoding, at best it won't work in some cases, at worst you are introducing security vulnerabilities. You can try, but there is just no way to do it right.

http://michaelthelin.se/security/2014/06/08/web-security-cro...

tialaramex5y ago

Generally you can safely treat text in an unknown encoding as UTF-8. Since you're expecting potential failures but want to press on anyway instead of causing an exception/ error you treat invalid sequences as U+FFFD the Replacement Character as you would in a language or API with no exception reporting mechanism.

There are lots of pleasing aspects to this choice. It's ASCII compatible of course, so anything that was actually ASCII is still ASCII, anything that was almost ASCII is just ASCII with U+FFFD where it deviated.

The replacement character resolutely isn't any of the specific things, nor any of the generic classes of thing you might be expected to treat differently for security reasons. It isn't a number, or a letter (of either "case"), it isn't white space, and it certainly isn't any of the separators, escapes or quote markers like ? or \ or + or . or _ or...

... yet it is still inside the BMP so it won't trigger weird (perhaps less well tested) behaviour for other planes.

It's self-synchronising. If something goes wrong somehow, in a few bytes if there is UTF-8 or an ASCII-compatible encoding the decoder will synchronise properly, you never end up "out of phase" as can happen for some encodings.

Most usefully, whatever you're now butted up against works with UTF-8 now. Maybe some day that'll get formally documented, maybe it won't. As the years drag on the chance of specifying _anything else_ shrink more, and the de facto popularity of UTF-8 means even if it's never formalised anywhere everybody will just assume UTF-8 anyway and you haven't to lift a finger.

2 more replies

at_a_remove5y ago

Oh, that's what I did of course. Some substitution, some Pokemon Exception Handling. At the end of the day, it was analysis, not a random file, so I wasn't worried about security.

What I am pointing out is that "know" is just doing a lot of magical work in that sentence.

BiteCode_dev5y ago

In a sense, it is simple: simple to understand.

Not simple to solve.

Like to win a race Usain Bolt, it's simple, run faster!

Fortunately, Python is well equipped for that. If you open a file with Python that you know it might contains mixed encoded text, you can use try/except to inform the user or open in binary mode, and just store the binary.

But my favorite way of doing it is:

    open('file', error=strategy)

Strategy can be:

- "ignore": undecodable text is skipped

- "replace": undecodable text is replace with "?"

- "surrogateescape" (you need to use utf8): undecodable text is decoded to a special representation which makes no human sense, but can en rencoded back to it's original value.

It's kinda ironic because people bashed Python for separating bytes/text, forcing them to deal with encoding correctly in Python 3. After all, this problem of "slamming bytes together" comes from languages that treat text as a bytes array, allowing this stupid mistake.

at_a_remove5y ago

Yeah, that was more or less my approach to analysis for these files. Essentially it was an export that was quasi-aware of Unicode and dumped out certain fields (of course they had to be variable length) in their original encodings, whatever they were. I got more than a few of these.

Unicode is great, as long as everyone upstream follows all of the rules and nothing goes wrong.

jbandela15y ago

Note: This post is basically a TLDR of https://www.theregister.com/2013/10/04/verity_stob_unicode/ by Verity Stob.

One of the reasons there is a lot of confusion about encodings vs Unicode is that Unicode was initially an encoding. It was thought that 65K characters was enough to represent all the characters in actual use across the languages and thus you just needed to change the from an 8 bit char to a 16 bit char and all would be well (apart from the issue of endianness). Thus Unicode initially specified what each symbol would look like encoded in 16bits. (see http://unicode.org/history/unicode88.pdf, particularly section 2). Windows NT, Java, ICU, all embraced this.

Then it turned out that you needed a lot more characters than 65K and instead of each character being 16 bits, you would need 32 bit characters (or else have weird 3 byte data types). Whereas people could justify going from 8 bits to 16 bits as a cost of not having to worry about charsets, most developers balked at 32 bits for every character. In addition, you now had a bunch of the early adopters (Java and Windows NT) that had already embraced 16 bit characters. So then encodings were hacked on such as UTF-16 (surrogate pairs of 16 bit characters for some unicode code points).

I think, if the problem had been understood better at the start that you have a lot more characters than will fit in 16 bits, then something UTF-8 would likely have been chosen as the canonical encoding and we could have avoided a lot of these issues. Alas, such is the benefit of 20/20 hindsight.

naniwaduni5y ago

> see http://unicode.org/history/unicode88.pdf, particularly section 2

it's Fascinating to see how people can arrive at the answer No and conclude that the answer is Yes

sgopalra5y ago

Interesting article from Joel spoolsky on unicode and character sets. https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...

dang5y ago

If curious see also

2015 https://news.ycombinator.com/item?id=9788253

2012: https://news.ycombinator.com/item?id=4771987

UpdatedFolders5y ago

I personally had a good time re-reading this over and over again when I was migrating python 2 to python 3, it's a great resource: http://farmdev.com/talks/unicode/

neiescOP5y ago

I think not explorer BOM UTF-8 https://en.wikipedia.org/wiki/Byte_order_mark

ExtremisAndy5y ago

I love C++ so much, and it has brought me such joy as a hobbyist programmer, but good grief, this one aspect of it (dealing with encodings & charsets) is so depressing I just want to cry sometimes.

nunez5y ago

F for respects for everyone who got wrecked by BOM (byte-order mark) and CRLF vs LF.

j / k navigate · click thread line to collapse

22 comments

at_a_remove5y ago

I was looking for the catch. Here it is: "It's really simple: Know what encoding a certain piece of text, that is, a certain byte sequence, is in, then interpret it with that encoding."

That's like "knowing" the truth. How?

Everything is simple -- until it isn't.

banthar5y ago

There is nothing you can do with text file with unknown encoding but treat it as an array of bytes.

If you start guessing the encoding, at best it won't work in some cases, at worst you are introducing security vulnerabilities. You can try, but there is just no way to do it right.

http://michaelthelin.se/security/2014/06/08/web-security-cro...

tialaramex5y ago

... yet it is still inside the BMP so it won't trigger weird (perhaps less well tested) behaviour for other planes.

2 more replies

at_a_remove5y ago

Oh, that's what I did of course. Some substitution, some Pokemon Exception Handling. At the end of the day, it was analysis, not a random file, so I wasn't worried about security.

What I am pointing out is that "know" is just doing a lot of magical work in that sentence.

BiteCode_dev5y ago

In a sense, it is simple: simple to understand.

Not simple to solve.

Like to win a race Usain Bolt, it's simple, run faster!

But my favorite way of doing it is:

    open('file', error=strategy)

Strategy can be:

- "ignore": undecodable text is skipped

- "replace": undecodable text is replace with "?"

- "surrogateescape" (you need to use utf8): undecodable text is decoded to a special representation which makes no human sense, but can en rencoded back to it's original value.

at_a_remove5y ago

Unicode is great, as long as everyone upstream follows all of the rules and nothing goes wrong.

jbandela15y ago

Note: This post is basically a TLDR of https://www.theregister.com/2013/10/04/verity_stob_unicode/ by Verity Stob.

naniwaduni5y ago

> see http://unicode.org/history/unicode88.pdf, particularly section 2

it's Fascinating to see how people can arrive at the answer No and conclude that the answer is Yes

sgopalra5y ago

Interesting article from Joel spoolsky on unicode and character sets. https://www.joelonsoftware.com/2003/10/08/the-absolute-minim...

dang5y ago

If curious see also

2015 https://news.ycombinator.com/item?id=9788253

2012: https://news.ycombinator.com/item?id=4771987

UpdatedFolders5y ago

I personally had a good time re-reading this over and over again when I was migrating python 2 to python 3, it's a great resource: http://farmdev.com/talks/unicode/

neiescOP5y ago

I think not explorer BOM UTF-8 https://en.wikipedia.org/wiki/Byte_order_mark

ExtremisAndy5y ago

I love C++ so much, and it has brought me such joy as a hobbyist programmer, but good grief, this one aspect of it (dealing with encodings & charsets) is so depressing I just want to cry sometimes.

nunez5y ago

F for respects for everyone who got wrecked by BOM (byte-order mark) and CRLF vs LF.

j / k navigate · click thread line to collapse