feat: add \U style 8-hex-digit unicode escape support in string literals

Extend the compiler's string parser to handle `\U` escapes with 8 hex digits
for codepoints above U+FFFF, replacing the previous TODO comment. The
`readUnicodeEscape` helper now accepts a variable digit count (4 or 8) and
the `readString` function dispatches `\u` with length 4 and `\U` with length 8.
New test cases verify correct encoding of emoji and ancient scripts, plus
error handling for incomplete long escapes and values that fit in 4 digits.
This commit is contained in:
Will Speak
2015-10-03 16:28:25 +00:00
parent 0f932e1f3a
commit 645296cbd8
4 changed files with 17 additions and 6 deletions
@@ -0,0 +1,2 @@
// expect error line 2
"\U01F603"
+8 -1
View File
@@ -14,4 +14,11 @@ System.print("\u0b83") // expect: ஃ
System.print("\u00B6") // expect: ¶
System.print("\u00DE") // expect: Þ
// TODO: Syntax for Unicode escapes > 0xffff?
// Big escapes:
var smile = "\U0001F603"
var byteSmile = "\xf0\x9f\x98\x83"
System.print(byteSmile == smile) // expect: true
System.print("<\U0001F64A>") // expect: <🙊>
System.print("<\U0001F680>") // expect: <🚀>
System.print("<\U00010318>") // expect: <𐌘>
@@ -0,0 +1,2 @@
// expect error line 2
"\U0060"