Skip to content

feat: authenticated file access (drive=), silent-HTML guard, GDFiles - #8

Merged
thorwhalen merged 2 commits into
masterfrom
claude/authenticated-file-access
Aug 19, 2026
Merged

thorwhalen merged 2 commits into
masterfrom
claude/authenticated-file-access

Conversation

@thorwhalen

Copy link
Copy Markdown
Member

Closes #5, closes #6, closes #7.

Motivated by Trufflepig-Travel/snout, which ingests client .xlsx exports from a private
Drive folder shared with the user's account. Today that is impossible: get_bytes is
unauthenticated and returns Google's sign-in page as if it were the file.

1. Authenticated single-file fetch (#5)

get_bytes(url, *, local_path=False, use_cache=False, drive=None, allow_html=False)

With drive=, the download goes through the API — CreateFile({'id': ...}) →
GetContentIOBuffer(), streamed and joined. Byte-exact, no temp file, no str round trip.
local_path and use_cache behave identically on both paths. drive=None keeps today's
public behaviour exactly.
get_bytes now also accepts a bare file id.

2. The silent-HTML failure (#6)

Drive serves its sign-in interstitial with HTTP 200, so the existing status check never
fired and the login page was returned as file content — plausible-looking bytes that only blow
up much later, in whatever tries to parse them. (Observed: a private 18.5 MB xlsx came back as
902 KB of HTML, no exception.)

_looks_like_html now sniffs the declared content type and the body prefix (tolerating
leading whitespace and a BOM, case-insensitive), and the public path raises
NotPubliclyShared — a dedicated, catchable exception whose message names the drive= fix and
the client_email sharing step. allow_html=True opts out when the file genuinely is HTML.

3. GDFiles (#7)

GDFiles(drive, *, folder_url=None, max_levels=None, include_hidden=False)

File-level Mapping keyed by file id or Drive URL (both normalised through
_extract_file_id, so they are one entry); values are bytes over the authenticated API.

__iter__ / __len__ require a scope. Unscoped, the mapping addresses the whole Drive —
unbounded and paginated, never what a caller wants, so offering it would be a trap. They raise
NotImplementedError with a message naming folder_url=; with a folder_url they yield that
folder's file ids using the same traversal GDReader uses. Lookup works either way, which is
the primary use case and needs no listing.

__contains__ answers from a metadata probe, never a download. A genuine 404 becomes KeyError
(Mapping contract), while 5xx / auth / network failures propagate as themselves rather than
being disguised as "absent".

4. Metadata without downloading

get_metadata(url_or_id, *, drive, fields=DEFAULT_METADATA_FIELDS) -> dict
GDFiles.metadata(key, *, fields=DEFAULT_METADATA_FIELDS) -> dict

title / fileSize / mimeType / modifiedDate — enough to decide whether an 18 MB fetch is
worth making. fileSize is coerced to int (Drive sends a string) and is absent for
Google-native files, which have no stored byte size.

Drive-by fixes

  • GDReader.__getitem__ corrupted binaries. It used
    GetContentString(mimetype='application/octet-stream').encode('latin-1'); GetContentString
    decodes as utf-8, so for any real binary this raised UnicodeDecodeError or silently
    mangled the bytes — latin-1 cannot undo a utf-8 decode. Now shares _download_via_api.
  • Folder traversal extracted to _iter_folder_files, shared by GDFiles and GDReader
    instead of duplicated. Semantics preserved exactly (regression-tested).
  • .gitignore now covers service-account keys, client_secrets.json, tokens, settings.yaml.
    Nothing sensitive was tracked before or is tracked now.
  • README: private-file quickstart, GDFiles, NotPubliclyShared troubleshooting, and the
    full service-account setup — including the step people forget, sharing the file/folder with
    the service account's client_email.

Tests

48 new tests in pydrivedol/tests/test_authenticated_access.py, all offline — no
credentials, no network. A FakeDrive implements the slice of the PyDrive2 API this package
calls; a FakeSession stands in for requests.Session. Covered: HTML detection (including
false-negative and false-positive cases, and a realistic sign-in-page fixture), the
authenticated download with local_path/use_cache, drive=None never touching the API,
metadata field selection and int coercion, all of GDFiles, and a regression test that
GDReader.__getitem__ returns exact bytes.

Full suite: 53 passed, 14 skipped (the skips are the pre-existing live-Drive tests).
ruff check and ruff format --check clean.

Not verified

No Google credentials available in this environment, so nothing was exercised against a real
Drive. Specifically unverified end to end: that GetContentIOBuffer() behaves as assumed
against the live API, and the exact Drive-v2 field names in DEFAULT_METADATA_FIELDS.

Follow-up (not in this PR)

[tool.pytest.ini_options] testpaths = ["pydrivedol"] excludes the repo-root tests/
directory, so tests/test_convert.py (4 tests) never runs in CI. They pass when invoked
directly. Left alone here to keep this PR to one concern.

https://claude.ai/code/session_01Vb8moaBd5Yg8b4Abo77nHf

Three related gaps that together made private-Drive ingest impossible.

Authenticated single-file fetch (#5)
  get_bytes gains a keyword-only drive=. With it, the download goes through
  the authenticated API (_download_via_api -> GetContentIOBuffer, streamed and
  joined; byte-exact, no temp file). Without it, behaviour is unchanged.
  get_bytes also now accepts a bare file id, not just a URL.

Silent-HTML failure (#6)
  Drive serves its sign-in interstitial with HTTP 200, so the public path was
  returning the login page as if it were the file -- plausible-looking bytes
  that only fail later, in whatever parses them. _looks_like_html now sniffs
  the content type and body prefix, and NotPubliclyShared is raised with an
  actionable message. allow_html=True opts out for genuine HTML downloads.

GDFiles (#7)
  A file-level Mapping keyed by file id or Drive URL (normalised through
  _extract_file_id), values bytes over the authenticated API. Iteration
  requires a folder_url scope -- listing a whole Drive is unbounded and
  paginated, so __iter__/__len__ raise NotImplementedError naming the fix
  rather than silently paginating. Lookup works either way. __contains__
  answers from a metadata probe, never a download; a genuine 404 becomes
  KeyError while 5xx/auth failures propagate as themselves.

Metadata without downloading
  get_metadata(url, drive=...) plus GDFiles.metadata: title/fileSize/mimeType/
  modifiedDate, so a caller can decide whether an 18MB fetch is worth it.
  fileSize is coerced to int; absent for Google-native files.

Also
  - GDReader.__getitem__ went through GetContentString(...).encode('latin-1'),
    which decodes as utf-8 first and so corrupted (or raised on) every real
    binary. It now shares _download_via_api.
  - Folder traversal extracted to _iter_folder_files, shared by GDFiles and
    GDReader instead of duplicated.
  - .gitignore covers service-account keys, client_secrets.json, tokens.
  - README: private-file quickstart, GDFiles, and the service-account setup
    steps (including sharing the folder with the account's client_email).

48 new tests, all offline: a fake GoogleDrive and a fake requests.Session
cover the HTML guard, the authenticated paths, metadata, and GDFiles.

Closes #5
Closes #6
Closes #7

Claude-Session: https://claude.ai/code/session_01Vb8moaBd5Yg8b4Abo77nHf
GDReader builds keys with os.path.join, so the nested-file key uses a
backslash on Windows. The Windows CI job is continue-on-error, but a
red check is still noise.

Claude-Session: https://claude.ai/code/session_01Vb8moaBd5Yg8b4Abo77nHf
@thorwhalen
thorwhalen merged commit c7f2b22 into master Aug 19, 2026
12 checks passed
@thorwhalen
thorwhalen deleted the claude/authenticated-file-access branch August 19, 2026 11:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant