Bind SQL keys as parameters; allowlist raw-SQL table names (fixes injection) - #7
Merged
Merged
Conversation
…names SqlBaseKvStore._mk_column_filter built the WHERE clause of __setitem__ and __delitem__ by formatting keys (and, for mapping keys, column names) into text(). A key containing a quote broke the write, and a key could change which rows the statement touched. - The key filter is now built from column expressions (table.c[col] == value), so values are bound parameters and column names can only be columns the table has (clear KeyError otherwise). This is how __getitem__ already matched keys, so reads and writes now agree. Empty mapping keys raise ValueError. - The legacy raw-SQL paths in sql_base (SqlTableRowsCollection, iter_rows) validate table names against an identifier allowlist and coerce LIMIT/OFFSET to integers. - Regression tests write, read back, overwrite and delete keys holding quotes, semicolons and --, and check no other row changes. Closes #6 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- Unknown column in a mapping key raises ValueError, not KeyError, so a typo isn't read as "key absent"; a case-insensitive unique match resolves, as unquoted SQL names did before. - Tuple/list keys raise a clear TypeError (single key column only). - Identifier allowlist accepts Unicode letters and digit-leading names (e.g. MySQL's 2020_sales), still rejecting all-digit names and anything that can end an identifier. - Drop two vacuous test assertions; add iter_rows table-name test. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #6. Follows up item 2 of the post-merge review on #4.
What was wrong
SqlBaseKvStore._mk_column_filter(used by__setitem__/__delitem__, so bySqlDictStore) formatted keys, and for mapping keys also column names, intotext(). A key containing'made writes fail, and a key could widen theWHEREclause: in a local SQLite reproduction, deleting a crafted, absent key removed every row. The legacy raw-SQL paths insql_base.pywrote table names into SQL text unchecked.Change
table.c[key_column] == key, or for a mapping keyand_(table.c[col] == val, ...). That is how__getitem__already matched keys, so reads and writes now compare keys the same way. Column names go throughtable.c, so only real columns can be named (SQLAlchemy quotes them); an unknown column raises aValueErrorthat lists the table's columns (notKeyError, so a typo isn't read as "key absent"). A name matching exactly one column case-insensitively still resolves, as unquoted SQL names did.WHEREand delete every row; it now raisesValueError. Tuple/list keys (which used to zip the characters of the column name into garbage SQL) raise a clearTypeError.None,float,UUID,datetimekeys, which raisedTypeErrorbefore, now compare like__getitem__does (None→IS NULL).SqlTableRowsCollectionanditer_rowsvalidate table names withvalidate_sql_identifier(Unicode letters, digits,_,$, not all digits, optionalschema.table; informativeValueErrorotherwise). Table names can't be bound, and these paths may run on a plain DB-API connection with no quoting helper, which is why this uses an allowlist and not quoting.LIMIT/OFFSETvalues are coerced withoperator.index.textimport dropped frombase.py.Tests
sqldol/tests/test_sql_injection.py: for 9 keys holding quotes, semicolons and--, insert, read back, overwrite, delete present, delete absent, and use as mapping-key values, each time checking the full table contents, so no other row may change. Plus normal str/int/dict keys, a test shaped like the known dependent's integer mapping-key delete, identifier allow/deny cases, and the legacy paths on a rawsqlite3connection.On the old code most of the key-path tests fail (the reviewer counted 38). With the fix, the whole suite passes locally (102 passed, doctests included, py3.10).
Review
An independent refute-review agent found no blockers. I applied its should-fix items in the second commit:
ValueErrorfor an unknown column, case-insensitive column resolution, a looser allowlist for MySQL-style digit-leading and Unicode names, two vacuous test assertions removed, and aTypeErrorfor sequence keys. Not changed:__getitem__still treats a mapping key as a key-column value while writes treat it as column conditions. That was already the case before this PR.Dependents
fleet_dependentslists one dependent, raglab_app. Its call sites usestrkeys (token, name, guid),intkeys (id, app_id) and__delitem__({"app_id": int, "user_id": int})on Postgres. Every column it names exists in the reflected table, and each key type binds to the same comparison the old literal SQL made. The new testtest_mapping_key_of_ints_on_integer_columnsmirrors the mapping-key call. raglab_app's own tests don't exercise sqldol (its paths need a live Postgres).Not changed (out of scope, noted in #6)
SQLAlchemyPersister.table_columns(DESCRIBE {self.table}) formats the ORM class, not caller input, and can't execute under SQLAlchemy 2.x.SqlDbCollection.from_config_dictformats a connection URI, not a query.sqlite3connection,SqlTableRowsCollection.count_rowscalls.first()(a SQLAlchemy result method), anditer_rowsnever detects the end of the table because a SELECT'srowcountis -1.🤖 Generated with Claude Code