Describe the bug
When a file larger than 4 MB is opened in the Query Tool, and a byte that isn't valid UTF-8 appears after the first 4 MB (for example, a Latin-1 é in an otherwise UTF-8 SQL dump), the file's first 4 MB are sent to the editor twice.
load_file streams read_file_generator(file_path, enc). enc comes from check_file_for_bom_and_binary(), which only looks at the first 1024 bytes, so it is utf-8 for such a file. read_file_generator() then reads the file in 4 MB chunks and yields each chunk as it goes. When a later chunk raises UnicodeDecodeError, the except branch reopens the file with latin-1 and yields it again from the start, after the earlier chunks have already been sent. If the user then saves the file, the duplicated content is written back.
To Reproduce
- Create a file of just over 5 MB that is UTF-8 except for one Latin-1 byte near the end:
data = b"SELECT 1;\r\n" * (5 * 1024 * 1024 // 11) + b"-- caf\xe9\r\n"
open("big.sql", "wb").write(data)
- Read it the way
load_file does:
from pgadmin.misc.file_manager import read_file_generator
out = "".join(read_file_generator("big.sql", "utf-8"))
print(len(data), len(out)) # 5242884 9437188
The output is 9,437,188 characters for a 5,242,884-byte file: the first 4 MB appear twice.
I checked this by calling the function on current master; I haven't gone through the UI with a file this size.
Expected behavior
The file's content appears exactly once, decoded as Latin-1 when it isn't valid in the detected encoding, which is what the fallback is meant to do.
Desktop
- pgAdmin version: current
master
- Mode: any (the file is read on the server)
Additional context
The same function uses codecs.open(), which is deprecated as of Python 3.14.
Would you like me to open a PR that fixes both? It would check which encoding applies before yielding anything, then read the file with the built-in open() (with newline='', so line endings stay exactly as they are).
Describe the bug
When a file larger than 4 MB is opened in the Query Tool, and a byte that isn't valid UTF-8 appears after the first 4 MB (for example, a Latin-1
éin an otherwise UTF-8 SQL dump), the file's first 4 MB are sent to the editor twice.load_filestreamsread_file_generator(file_path, enc).enccomes fromcheck_file_for_bom_and_binary(), which only looks at the first 1024 bytes, so it isutf-8for such a file.read_file_generator()then reads the file in 4 MB chunks and yields each chunk as it goes. When a later chunk raisesUnicodeDecodeError, theexceptbranch reopens the file withlatin-1and yields it again from the start, after the earlier chunks have already been sent. If the user then saves the file, the duplicated content is written back.To Reproduce
load_filedoes:I checked this by calling the function on current
master; I haven't gone through the UI with a file this size.Expected behavior
The file's content appears exactly once, decoded as Latin-1 when it isn't valid in the detected encoding, which is what the fallback is meant to do.
Desktop
masterAdditional context
The same function uses
codecs.open(), which is deprecated as of Python 3.14.Would you like me to open a PR that fixes both? It would check which encoding applies before yielding anything, then read the file with the built-in
open()(withnewline='', so line endings stay exactly as they are).