fix: standardize open() calls with explicit encoding='utf-8' across codebase - #7366
ANIRUDDHA ADAK (aniruddhaadak80) wants to merge 1 commit into
Conversation
|
ANIRUDDHA ADAK (@aniruddhaadak80) please read the following Contributor License Agreement(CLA). If you agree with the CLA, please reply with the following information.
Contributor License AgreementContribution License AgreementThis Contribution License Agreement (“Agreement”) is agreed to by the party signing below (“You”),
|
There was a problem hiding this comment.
Pull request overview
Repository-wide hardening against Windows locale-dependent decoding by standardizing text file I/O to use explicit UTF-8 encoding (addresses UnicodeDecodeError class of failures described in #5566).
Changes:
- Add
encoding="utf-8"toopen()calls across samples, packages, docs tooling, and tests. - Add
encoding="utf-8"toaiofiles.open()calls in FastAPI samples. - Minor whitespace-only adjustments in a couple of sample files included in the diff hunks.
Reviewed changes
Copilot reviewed 51 out of 51 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| python/samples/task_centric_memory/utils.py | Use UTF-8 when reading YAML. |
| python/samples/gitty/src/gitty/_gitty.py | Use UTF-8 when reading global/local config files. |
| python/samples/gitty/src/gitty/main.py | Use UTF-8 when creating the config file. |
| python/samples/core_streaming_response_fastapi/app.py | Use UTF-8 when reading model config via aiofiles. |
| python/samples/core_streaming_handoffs_fastapi/app.py | Use UTF-8 for model config and chat history file reads via aiofiles. |
| python/samples/core_streaming_handoffs_fastapi/agent_user.py | Use UTF-8 when writing chat history JSON. |
| python/samples/core_distributed-group-chat/_utils.py | Use UTF-8 when reading YAML config. |
| python/samples/core_chess_game/main.py | Use UTF-8 when reading model config. |
| python/samples/core_chainlit/app_team.py | Use UTF-8 when reading model config (plus whitespace-only change). |
| python/samples/core_chainlit/app_agent.py | Use UTF-8 when reading model config (plus whitespace-only change). |
| python/samples/agentchat_streamlit/agent.py | Use UTF-8 when reading model config. |
| python/samples/agentchat_fastapi/app_team.py | Use UTF-8 for config/state/history reads and writes via aiofiles. |
| python/samples/agentchat_fastapi/app_agent.py | Use UTF-8 for config/state/history reads and writes via aiofiles. |
| python/samples/agentchat_chess_game/main.py | Use UTF-8 when reading model config. |
| python/samples/agentchat_chainlit/app_team_user_proxy.py | Use UTF-8 when reading model config. |
| python/samples/agentchat_chainlit/app_team.py | Use UTF-8 when reading model config. |
| python/samples/agentchat_chainlit/app_agent.py | Use UTF-8 when reading model config. |
| python/packages/magentic-one-cli/src/magentic_one_cli/_m1.py | Use UTF-8 when reading the default CLI config. |
| python/packages/autogen-studio/tests/test_team_manager.py | Use UTF-8 when writing temporary JSON/YAML configs in tests. |
| python/packages/autogen-studio/tests/test_lite_studio.py | Use UTF-8 when reading/writing JSON and env files in tests. |
| python/packages/autogen-studio/autogenstudio/web/auth/manager.py | Use UTF-8 when reading YAML auth config. |
| python/packages/autogen-studio/autogenstudio/lite/studio.py | Use UTF-8 when writing the generated env file. |
| python/packages/autogen-studio/autogenstudio/gallery/builder.py | Use UTF-8 when writing default gallery JSON. |
| python/packages/autogen-studio/autogenstudio/database/schema_manager.py | Use UTF-8 when writing Alembic files (plus flagged env.py template issue). |
| python/packages/autogen-studio/autogenstudio/cli.py | Use UTF-8 when writing the temporary env file. |
| python/packages/autogen-ext/tests/test_filesurfer_agent.py | Use UTF-8 when writing a temporary HTML test file via aiofiles. |
| python/packages/autogen-ext/tests/task_centric_memory/utils.py | Use UTF-8 when reading YAML in tests. |
| python/packages/autogen-ext/tests/code_executors/test_user_defined_functions.py | Ensure embedded code blocks open files with UTF-8. |
| python/packages/autogen-ext/tests/code_executors/test_docker_jupyter_code_executor.py | Ensure embedded code blocks open files with UTF-8. |
| python/packages/autogen-ext/tests/code_executors/test_docker_commandline_code_executor.py | Ensure embedded code blocks open files with UTF-8. |
| python/packages/autogen-ext/tests/code_executors/test_commandline_code_executor.py | Ensure embedded code blocks open files with UTF-8. |
| python/packages/autogen-ext/tests/code_executors/test_aca_dynamic_sessions.py | Ensure embedded code blocks open files with UTF-8. |
| python/packages/autogen-ext/src/autogen_ext/experimental/task_centric_memory/utils/page_logger.py | Use UTF-8 when writing HTML/log artifacts. |
| python/packages/autogen-ext/src/autogen_ext/experimental/task_centric_memory/utils/chat_completion_client_recorder.py | Use UTF-8 when reading/writing session recordings. |
| python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_docker_jupyter.py | Use UTF-8 when writing HTML output artifacts. |
| python/packages/autogen-ext/examples/mcp_session_host_example.py | Use UTF-8 when reading JSON/YAML model config. |
| python/packages/agbench/src/agbench/tabulate_cmd.py | Use UTF-8 when reading console logs. |
| python/packages/agbench/src/agbench/run_cmd.py | Use UTF-8 when reading scenario/env/config files and writing scripts. |
| python/packages/agbench/src/agbench/remove_missing_cmd.py | Use UTF-8 when reading console logs. |
| python/packages/agbench/src/agbench/linter/coders/oai_coder.py | Use UTF-8 for cache reads/writes. |
| python/packages/agbench/src/agbench/linter/cli.py | Use UTF-8 when reading log files. |
| python/packages/agbench/benchmarks/process_logs.py | Use UTF-8 when reading benchmark logs/prompts/messages. |
| python/packages/agbench/benchmarks/HumanEval/Templates/AgentChat/scenario.py | Use UTF-8 when reading config/prompt. |
| python/packages/agbench/benchmarks/HumanEval/Templates/AgentChat/custom_code_executor.py | Use UTF-8 when reading test code. |
| python/packages/agbench/benchmarks/GAIA/Templates/SelectorGroupChat/scenario.py | Use UTF-8 when reading config/prompt. |
| python/packages/agbench/benchmarks/GAIA/Templates/ParallelAgents/scenario.py | Use UTF-8 when reading config/prompt and writing per-team console logs. |
| python/packages/agbench/benchmarks/GAIA/Templates/MagenticOne/scenario.py | Use UTF-8 when reading config/prompt. |
| python/packages/agbench/benchmarks/GAIA/Scripts/custom_tabulate.py | Use UTF-8 when reading expected answers and console logs. |
| python/fixup_generated_files.py | Use UTF-8 when reading/writing files during fixup. |
| python/docs/src/generate_api_reference.py | Use UTF-8 when writing generated docs files. |
| python/docs/redirects/redirects.py | Use UTF-8 when reading redirects config and writing redirect pages (plus whitespace-only change). |
| run_migrations_offline() | ||
| else: | ||
| run_migrations_online()""" | ||
|
|
||
| with open(env_path, "w") as f: | ||
| with open(env_path, "w", encoding="utf-8") as f: |
There was a problem hiding this comment.
The _create_minimal_env_py template string appears to generate invalid Python: inside context.configure(...), compare_type=True is missing a trailing comma (both in offline and online sections). This will produce a syntax error in the generated env.py and break Alembic migrations. Add the missing commas in the template content.
|
Copilot open a new pull request to apply changes based on the comments in this thread |
…' across codebase As raised in microsoft#5566, opening files without specifying encoding causes UnicodeDecodeErrors in non-English environments on Windows, particularly environments handling CJK (cp950) encodings. Fixing this codebase-wide. Resolves microsoft#5566 (indirectly by expanding fix to all files).
4e20166 to
d4477e4
Compare
Related to issue #5566. Calling
open()without explicitly passingencoding='utf-8'causes fatalUnicodeDecodeErrors on Windows machines operating in non-English environments (such as cp1252 or cp950).While the root issue reported in #5566 was fixed for one file, this PR acts as a comprehensive repository-wide cleanup, adding explicit utf-8 encodings to over 60 Python files where text paths were opened without it.