Conversation
The string rules excluded every eol character, so loads() rejected the two separators that JSON and the JSON5 spec allow in strings, e.g. the output of JSON.stringify for such a string.
|
Interesting bug! I'm curious, how did you find this? I think the fix looks correct, but this isn't the way I'd prefer it be written. I'd argue that the problem is that the line in the grammar So, what you probably really want is
and then the rule would be (and you'd change the Although, from a performance point of view, given the current implementation, it'd probably be slightly faster to have the line be: Since the common case is actually that rule, and not a backslash followed by something. From your comment, it sounds like you actually figured out how to use Given that, do you want to try to rewrite the fix the way I'm suggesting? I haven't actually coded this myself, so I'm not 100% sure I'm right, and I'd want to code it and run the tests to see :). Separately, in the tests, please don't embed |
json5.loadsrejects strings containing an unescaped U+2028 or U+2029, such as the output of JavaScript'sJSON.stringify("a b"), whichjson.loadsaccepts. The JSON5 spec lists both code points as valid string characters ("Like JSON, JSON5 allows the Unicode code points U+2028 and U+2029 to appear unescaped in strings"), and the reference json5 parser accepts them.The string rules excluded every
eolcharacter; this adds the two code points as alternatives injson5.gand regeneratesparser.py. The current glop doesn't emit the hand-addedposargument inParser.__init__, so I kept that hunk as it was.The tests cover both quote styles, a backslash before U+2028 still being a line continuation, and a raw CR still failing. Compared with json5 2.2.3 on json5-tests and about 33k generated strings, the only results that change are these cases, and all of them now match the reference.
python run testsandpython run checkpass.Written with AI assistance (Claude); I have reviewed the change.