Conversation
…ash error labels chumsky 0.10 produces byte offsets over &str, while ariadne and the rest of the compiler read spans as char offsets. Convert lexer token spans and interpolated-string spans to char offsets. Closes #6377
# Conflicts: # CHANGELOG.md
# Conflicts: # CHANGELOG.md
# Conflicts: # CHANGELOG.md
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Since the chumsky 0.10 lexer migration (#5223), token spans have been byte offsets: chumsky's
SimpleSpancounts bytes over&str. Everything that reads spans counts chars: ariadne, the lexer's ownconvert_lexer_error, and the parser before #5223. So each multi-byte character shifts every later span to the right. A label lands on the wrong column, and once a span runs past the source's char length, theassert!inerror_message.rspanics. This PR converts token spans to char offsets once lexing finishes, using a byte→char table built once per source and skipped for ASCII input. It applies the same conversion to spans inside interpolated strings, which had the same bug.Why char offsets rather than switching everything to bytes. Chars are the convention the rest of the compiler already uses, and they're what spans held before #5223. So this restores the earlier contract for the bindings and
debug annotate, instead of moving every consumer to a new one.Tests
test_error_after_non_ascii(integration,error_messages.rs) covers three cases: the panic from the issue, the one-column label shift after# café, and an interpolation parse error afteréééé, which also panicked. The test fails onmainwithspan Some(1:67-78) is out of bounds of the source (len = 74).test_lex_source_non_ascii_spanspins char spans for tokens afteréandü.test_unicodesnapshot forfrom tètechanges from0-10to0-9, the query's actual char length. The old snapshot had captured the byte-offset bug.prqlc debug annotateby hand. After a line containing non-ASCII text, it now attaches frames to the right lines; onmainthey shift down one line.This PR leaves the
assert!inerror_message.rsunchanged. A different span bug could still reach it, so turning it into a location-less error remains a separate option.Closes #6377 — automated triage