diff options
| author | Jeff Epler <jepler@gmail.com> | 2020-09-08 20:54:47 -0500 |
|---|---|---|
| committer | Jeff Epler <jepler@gmail.com> | 2020-09-08 20:54:47 -0500 |
| commit | 0eee93729a3074461c4907f03deb803c2f35a6fe (patch) | |
| tree | a19db09e206d69ed4d409d9d81bba0b43787319e | |
| parent | bdb07adfccaf25576a8f1e074db7f8b7e48890b2 (diff) | |
Fix decompression of unicode values above 2047
Two problems: The lead byte for 3-byte sequences was wrong, and one
mid-byte was not even filled in due to a missing "++"!
Apparently this was broken ever since the first "Compress as unicode,
not bytes" commit, but I believed I'd "tested" it by running on the
Pinyin translation.
This rendered at least the Korean and Japanese translations completely
illegible, affecting 5.0 and all later releases.
| -rw-r--r-- | supervisor/shared/translate.c | 4 |
1 files changed, 2 insertions, 2 deletions
diff --git a/supervisor/shared/translate.c b/supervisor/shared/translate.c index 0235293be..5cd7b8dd8 100644 --- a/supervisor/shared/translate.c +++ b/supervisor/shared/translate.c @@ -59,8 +59,8 @@ STATIC int put_utf8(char *buf, int u) { *buf = 0b10000000 | (u & 0b00111111); return 2; } else { // u <= 0xffff - *buf++ = 0b11000000 | (u >> 12); - *buf = 0b10000000 | ((u >> 6) & 0b00111111); + *buf++ = 0b11100000 | (u >> 12); + *buf++ = 0b10000000 | ((u >> 6) & 0b00111111); *buf = 0b10000000 | (u & 0b00111111); return 3; } |
