Conversation
…upting the string
A \x followed by one hex digit dropped that digit ("\x4g" became "xg"), and
\u{...} accepted an empty, unclosed or out-of-range escape, emitting a NUL, the
code point without its brace, or U+FFFD. Both now rewind and keep the text as
written, as an unknown escape already does.
for more information, see https://pre-commit.ci
prql-bot
left a comment
There was a problem hiding this comment.
One behavior change in interpolated strings isn't covered by the description; see inline.
| c | ||
| } | ||
| None => { | ||
| input.rewind(checkpoint); |
There was a problem hiding this comment.
Rewinding here also affects interpolated strings, because interpolation() lexes f- and s-strings with escaping on and then parses {...} out of the result. On main, f"\u{x}" and s"\u{x}" fail to compile (expected '}', but found end of input). With this change, f"\u{x}" compiles to CONCAT('u', x) and s"\u{x}" splices ux into the SQL. A malformed escape turns into a silent interpolation of whatever name follows it. Going the other way, f"\u{41" used to give 'A' and now fails with expected '{' or interpolated string variable.
This matches what an unknown escape already does (f"\q{x}" interpolates x), so it may be acceptable. It isn't mentioned in the description or the changelog, though, and whether a malformed \u{ inside f-/s-strings should interpolate or error is a maintainer's call.
# Conflicts: # CHANGELOG.md
# Conflicts: # CHANGELOG.md
A string's
\xor\u{...}escape that isn't well-formed no longer changes the string's contents. Before this, the lexer quietly rewrote them:"\x4g"'xg'(the4is dropped)'x4g'"\u{}z"z'u{}z'"\u{41"(no closing brace)'A''u{41'"\u{110000}"(not a valid code point)U+FFFD'u{110000}'Well-formed escapes (
"\x41","\u{1F600}") are unchanged. A malformed escape is now handled like an unknown escape such as"\q": the backslash is dropped and the rest is kept as written. Inparse_escape_sequence, both arms now save the position before reading hex digits and rewind to it when the escape doesn't match the form the strings reference documents: exactly two digits for\xhh, and 1–6 digits, a closing}and a valid code point for\u{...}. Before, those arms kept whatever digits they had already consumed.Raising a lexer error would be stricter. I kept the lenient behavior because it matches how unknown escapes are already treated, and because maintainers have previously said permissive lexing is fine unless it causes a problem (#3476). Here the problem was data loss, which rewinding fixes.
Cases added to
lexer::test::quotes; the\x4gcase fails onmain(left: "xg",right: "x4g").cargo test -p prqlc-parserandcargo test -p prqlcpass locally.