Skip to content

fix: convert rows in batches for fetchone() and cursor iteration - #982

Open
maharanay22 wants to merge 1 commit into
databricks:mainfrom
maharanay22:fix-row-iteration-performance
Open

maharanay22 wants to merge 1 commit into
databricks:mainfrom
maharanay22:fix-row-iteration-performance

Conversation

@maharanay22

Copy link
Copy Markdown

What type of PR is this?

  • Bug Fix

Description

Fixes #981.

Iterating a cursor (for row in cursor) or calling fetchone() in a loop costs about 0.6 to 0.8 ms per row, more than 100x slower than fetchall(). fetchone() fetched a single row and passed it through _convert_arrow_table, which builds a pandas DataFrame on every call. SQLAlchemy's default Result iteration calls cursor.fetchone() per row, so a 1M-row read through SQLAlchemy takes many minutes instead of seconds

This PR moves fetchone(), fetchmany() and fetchall() into the ResultSet base class so Thrift, SEA and kernel share one implementation. Each backend now provides _fetchmany_table, _fetchall_table and _convert_table. fetchone() fetches and converts arraysize rows at a time and serves them from a small buffer that keeps both the converted rows and the raw batch. Every other fetch method (fetchmany, fetchall, fetchmany_arrow, fetchall_arrow, and the columnar/JSON variants) drains that buffer first, so ordering and counts stay exact when row and Arrow fetches are mixed, for example fetchone() followed by fetchall_arrow(). rownumber still reports the rows handed to the caller. The buffer holds at most one batch of arraysize rows and is released as soon as it is drained

Timings on 20,000 rows (bigint + string) from an in-memory Arrow queue:

fetchall for row in cursor fetchone loop
pandas 3.0.6, before 0.07s 16.1s 11.6s
pandas 3.0.6, after 0.10s 0.10s 0.14s
pandas 2.3.3, before 0.06s 11.9s 14.8s
pandas 2.3.3, after 0.13s 0.10s 0.13s

One trade-off: the first fetchone() now fills up to arraysize rows, so on Thrift it can make the next FetchResults call earlier than before when the direct results hold fewer rows than that. Lowering cursor.arraysize restores the old behavior

How is this tested?

  • Unit tests
  • Manually

New tests in test_fetches.py, test_sea_result_set.py and test_kernel_result_set.py count calls to the conversion function and check that it runs once per batch instead of once per row. These fail on main (21, 6 and 10 calls instead of 3, 2 and 2) and pass with this change. Other new tests cover the interleavings: fetchone() then fetchall(), iteration across several queue batches, fetchone() then fetchmany() then fetchall_arrow(), fetchmany_arrow() spanning buffered and new rows, ColumnQueue and JSON paths, empty results, and rownumber at each step. test_sea_result_set.py::test_fetchone now asserts on rownumber instead of the private _next_row_index. Full unit suite: 1045 passed, 5 skipped with pandas 3, and the touched suites pass with pandas 2.3. black --check src and mypy src pass. The timings above were measured manually on both pandas versions

Related Tickets & Documents

Fixes #981. Affects every backend's fetchone() and therefore SQLAlchemy result iteration

fetchone() fetched a single row and ran it through _convert_arrow_table,
which builds a pandas DataFrame per call. That fixed cost made iterating a
cursor, or any fetchone() loop such as SQLAlchemy's default result
iteration, about 1 ms per row, more than 100x slower than fetchall().

ResultSet now implements fetchone(), fetchmany() and fetchall() once for
all backends. fetchone() fetches and converts arraysize rows at a time and
serves them from a small buffer that keeps both the converted rows and the
raw batch. Every other fetch method drains that buffer first, so ordering
and counts are unchanged when row and Arrow/columnar/JSON fetches are
mixed, and rownumber still reports the number of rows handed to the
caller.

Signed-off-by: Maha Rana Yadavalli <271375718+maharanay22@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Iterating a cursor or calling fetchone() in a loop is 100x+ slower than fetchall() (affects SQLAlchemy)

1 participant