Skip to content

fix: count token spans in chars so non-ASCII text doesn't shift or crash error labels - #6378

Open
prql-bot wants to merge 7 commits into
mainfrom
fix/issue-6377
Open

prql-bot wants to merge 7 commits into
mainfrom
fix/issue-6377

Conversation

@prql-bot

Copy link
Copy Markdown
Collaborator

Since the chumsky 0.10 lexer migration (#5223), token spans have been byte offsets: chumsky's SimpleSpan counts bytes over &str. Everything that reads spans counts chars: ariadne, the lexer's own convert_lexer_error, and the parser before #5223. So each multi-byte character shifts every later span to the right. A label lands on the wrong column, and once a span runs past the source's char length, the assert! in error_message.rs panics. This PR converts token spans to char offsets once lexing finishes, using a byte→char table built once per source and skipped for ASCII input. It applies the same conversion to spans inside interpolated strings, which had the same bug.

Why char offsets rather than switching everything to bytes. Chars are the convention the rest of the compiler already uses, and they're what spans held before #5223. So this restores the earlier contract for the bindings and debug annotate, instead of moving every consumer to a new one.

Tests

  • test_error_after_non_ascii (integration, error_messages.rs) covers three cases: the panic from the issue, the one-column label shift after # café, and an interpolation parse error after éééé, which also panicked. The test fails on main with span Some(1:67-78) is out of bounds of the source (len = 74).
  • test_lex_source_non_ascii_spans pins char spans for tokens after é and ü.
  • The existing test_unicode snapshot for from tète changes from 0-10 to 0-9, the query's actual char length. The old snapshot had captured the byte-offset bug.
  • I also checked prqlc debug annotate by hand. After a line containing non-ASCII text, it now attaches frames to the right lines; on main they shift down one line.

This PR leaves the assert! in error_message.rs unchanged. A different span bug could still reach it, so turning it into a location-less error remains a separate option.

Closes #6377 — automated triage

…ash error labels

chumsky 0.10 produces byte offsets over &str, while ariadne and the rest of
the compiler read spans as char offsets. Convert lexer token spans and
interpolated-string spans to char offsets.

Closes #6377

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Non-ASCII text before an error panics prqlc or shifts the error label (byte vs char spans)

1 participant