Repository navigation
Strip leading UTF-8 BOM when reading INI files - #76
Pitchfork-and-Torch wants to merge 2 commits into
Conversation
Windows editors often write a UTF-8 BOM; with encoding=utf-8 that left U+FEFF on the first section line and raised ParseError. Strip the BOM after read for both IniConfig() and IniConfig.parse().
RonnyPfannschmidt
left a comment
There was a problem hiding this comment.
🤖 Written by Claude Opus 5.5 via Claude Code for the iniconfig maintainers; I prompted it, it did the work, I read it.
Thanks, stripping a leading BOM is a welcome fix. Please make three changes before this goes in:
- Strip it in one place. The same block is currently in both
IniConfig.__init__andIniConfig.parse. Move it into_parse.parse_ini_data, which every path (file read anddata=) already goes through. - Use a named constant. Define the BOM once at module level, e.g.
UTF8_BOM: Final = "", and use that name instead of an inline literal. - Write it as an escape. The current diff contains a raw U+FEFF character inside
"...", which is invisible in editors and in review. Spell it""so it is readable. Thestartswithcheck is then unnecessary:data = data.removeprefix(UTF8_BOM)does the same thing.
Separately, pre-commit.ci is red: mypy wants annotations on the three new tests (-> None and tmp_path: Path).
Generated by Claude Code
|
🤖 Written by Claude Opus 5.5 via Claude Code for the iniconfig maintainers; I prompted it, it did the work, I read it. Correction to my review: both snippets in it were mangled into the raw invisible BOM. The intended code is: UTF8_BOM: Final = "\N{BYTE ORDER MARK}"
data = data.removeprefix(UTF8_BOM)Generated by Claude Code |
Strip once in parse_ini_data, name the mark with \N{BYTE ORDER MARK}, and annotate the new tests.
|
Moved the strip into parse_ini_data. The constant is UTF8_BOM: Final = "\N{BYTE ORDER MARK}", and the three new tests now have return annotations plus a Path type on tmp_path. |
RonnyPfannschmidt
left a comment
There was a problem hiding this comment.
Well done thanks
do you want to squash or should I squash merge
|
Please squash merge — happy either way as long as it lands as one commit. Thanks! |
Summary
Windows editors often write a UTF-8 BOM. With the default
encoding=\"utf-8\", that left U+FEFF on the first section line and raisedParseError: unexpected line.Strip a leading BOM after read for both
IniConfig()andIniConfig.parse(), including when content is passed viadata=.Test plan
test_utf8_bom_file,test_utf8_bom_in_data_string,test_parse_utf8_bom_filepass