Skip to content

fix(sql_base): iter_rows detects the end of a table; count_rows works on DB-API - #9

Merged
thorwhalen merged 3 commits into
masterfrom
fix-legacy-sqlite-iter-rows
Sep 22, 2026
Merged

thorwhalen merged 3 commits into
masterfrom
fix-legacy-sqlite-iter-rows

Conversation

@thorwhalen

@thorwhalen thorwhalen commented Sep 22, 2026 •

Copy link
Copy Markdown
Member

Fixes the two pre-existing legacy-path bugs noted in the bodies of #7 and #8.

What was wrong

  • sql_base.iter_rows decided whether a page had rows from rowcount. For a SELECT on sqlite3 (and other DB-API drivers that don't pre-buffer), rowcount is -1, which is truthy. So without a small limit it kept requesting empty pages until limit (default 1e12) ran out. SqlTableRowsCollection.__iter__ goes through it, so iterating a collection on a sqlite3 connection never ended.
  • Also in iter_rows: the running counter restarted each page at the previous page's last index. When limit spanned several pages it yielded extra rows. For example, batch_size=2, limit=5 on a 7-row table gave 6 rows.
  • SqlTableRowsCollection.count_rows called .first(), a SQLAlchemy result method. On a DB-API connection, execute returns a cursor, which has no first, so len(collection) raised AttributeError.

Change

  • iter_rows stops on the first page shorter than it asked for, and never asks for more rows than remain under limit. Offset/limit act like a slice rows[offset:offset+limit]. A batch_size < 1 raises a ValueError that names it (it was a bare range() error before). Doctests added.
  • count_rows uses fetchone()[0], which a DB-API cursor has. (The raw-SQL strings in this module only execute on DB-API connections anyway: SQLAlchemy 2 rejects a bare string. That is unchanged and out of scope here.)
  • Arguments are validated when iter_rows is called, not on the first next(). A negative offset (sqlite reads it as 0 and duplicated rows) and a negative limit (silently []) now raise ValueError. SqlTableRowsCollection[a:0] returned every row (if stop:); it is now empty.
  • The bounded-batch workaround and its note in test_sql_injection.py are dropped, since the unbounded call now terminates.

Tests

sqldol/tests/test_legacy_raw_sql_rows.py, on in-memory sqlite3, uses a connection wrapper that counts queries and fails past a cap. That way the old code fails instead of hanging. It covers end detection for several batch sizes (with an exact query count), the empty table, 32 offset/limit/batch combinations checked against list slicing, count_rows/len/iteration/slicing on the collection, and SqlTableRowsSequence. Against master, 12 of the 41 fail. On this branch the full suite with doctests passes: 149 passed (py3.10).

Dependents

fleet_dependents lists raglab_app. It imports only sqldol.stores and sqldol.base, never sql_base, so no call site changes. Its own test suite cannot run here for reasons unrelated to sqldol: test_app.py imports a moved name, and the other tests call a LangChain method that no longer exists. sqldol.stores imports fine from its environment with this branch installed.

Review

An independent refute-review agent found no blockers. It swept old vs new iter_rows over batch_size × offset × limit and confirmed that the only differences on valid arguments are the fixed over-yield cases. I applied its should-fix items in the second commit: eager validation, rejecting negative offset/limit, and an accurate count_rows note. I also applied its rows[:0] nit. Not changed: paging without ORDER BY can be inconsistent on non-sqlite backends or on a table that changes mid-iteration. That was already true before this PR.

🤖 Generated with Claude Code

thorwhalen and others added 3 commits September 22, 2026 15:25
…-API

- iter_rows decided whether a page had rows from the cursor's rowcount,
  which is -1 (truthy) for a SELECT on sqlite3, so without a limit it
  requested empty pages until limit (default 1e12) ran out. It now stops
  on the first page shorter than requested. It also no longer yields more
  than limit rows when limit spans several pages (the running counter
  restarted at the last index), and a page never asks for more rows than
  remain under limit. A batch_size < 1 raises a ValueError naming it.
- SqlTableRowsCollection.count_rows used .first(), a SQLAlchemy result
  method a DB-API cursor lacks; it now uses fetchone(), which both have.

Tests: sqldol/tests/test_legacy_raw_sql_rows.py (sqlite3, with a
query-counting connection so the old code fails instead of hanging).
The workaround note in test_sql_injection.py is dropped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…count_rows note

- iter_rows validates batch_size/offset/limit when called (not on first
  next()); a negative offset duplicated rows on sqlite and a negative
  limit silently gave [], both now ValueError.
- SqlTableRowsCollection[a:0] returned every row (`if stop:`).
- count_rows comment no longer claims SQLAlchemy 2 connections work: the
  raw-SQL strings here only execute on DB-API connections.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@thorwhalen
thorwhalen merged commit d746553 into master Sep 22, 2026
6 checks passed
@thorwhalen
thorwhalen deleted the fix-legacy-sqlite-iter-rows branch September 22, 2026 15:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant