Skip to content

Query Tool: first 4 MB of a large file is loaded twice when a later chunk isn't valid UTF-8 #10462

Description

@Barry-Lenhart

Describe the bug

When a file larger than 4 MB is opened in the Query Tool, and a byte that isn't valid UTF-8 appears after the first 4 MB (for example, a Latin-1 é in an otherwise UTF-8 SQL dump), the file's first 4 MB are sent to the editor twice.

load_file streams read_file_generator(file_path, enc). enc comes from check_file_for_bom_and_binary(), which only looks at the first 1024 bytes, so it is utf-8 for such a file. read_file_generator() then reads the file in 4 MB chunks and yields each chunk as it goes. When a later chunk raises UnicodeDecodeError, the except branch reopens the file with latin-1 and yields it again from the start, after the earlier chunks have already been sent. If the user then saves the file, the duplicated content is written back.

To Reproduce

  1. Create a file of just over 5 MB that is UTF-8 except for one Latin-1 byte near the end:
    data = b"SELECT 1;\r\n" * (5 * 1024 * 1024 // 11) + b"-- caf\xe9\r\n"
    open("big.sql", "wb").write(data)
  2. Read it the way load_file does:
    from pgadmin.misc.file_manager import read_file_generator
    out = "".join(read_file_generator("big.sql", "utf-8"))
    print(len(data), len(out))  # 5242884 9437188
    The output is 9,437,188 characters for a 5,242,884-byte file: the first 4 MB appear twice.

I checked this by calling the function on current master; I haven't gone through the UI with a file this size.

Expected behavior

The file's content appears exactly once, decoded as Latin-1 when it isn't valid in the detected encoding, which is what the fallback is meant to do.

Desktop

  • pgAdmin version: current master
  • Mode: any (the file is read on the server)

Additional context

The same function uses codecs.open(), which is deprecated as of Python 3.14.

Would you like me to open a PR that fixes both? It would check which encoding applies before yielding anything, then read the file with the built-in open() (with newline='', so line endings stay exactly as they are).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions