Skip to content

Latest commit

 

History

History
1769 lines (1375 loc) · 86.1 KB

File metadata and controls

1769 lines (1375 loc) · 86.1 KB

SatQuery AI — Project Status

Hackathon: SIH 2026 Problem Statement: SatQuery AI — An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries Repo root: D:\SatQueryAI (recovered from F: drive exFAT failure; see section 10) Last updated: 2026-08-26 Current milestone: M3-D · checkpoint-1590 accepted as the final adapter; eval harness, unit tests and evidence doc complete; base-vs-adapted numbers pending a GPU session Milestone scheme: unified (section 4). Older numbering is mapped in 4.2 — do not reintroduce it.


1. Environment audit

1.1 Host machine

Item Value Verdict
OS Windows 11 Home Single Language, 10.0.26200, 64-bit OK
Shell PowerShell 7 (primary) + Git Bash OK
CPU Intel Core i5-10300H @ 2.50 GHz — 4 cores / 8 threads Adequate for orchestration, not for training
RAM 15.9 GB total, 3.2 GB free at audit time Tight — close browsers/IDEs before local runs
GPU NVIDIA GeForce GTX 1650, 4 GB VRAM, driver 576.52, CUDA 12.9 Too small to train or host a 3B VLM
iGPU Intel UHD Graphics (1 GB) Not usable for compute
Disk C: 30.7 GB free / D: 289.3 GB / E: 305.4 GB / F: 797.2 GB free (project drive) Far above the 100 GB budget

Conclusion: the laptop is an orchestration and UI machine. All training and heavy inference goes to Colab GPU. The GTX 1650 is kept only as a demo-safety fallback for small CPU/GPU-lite paths (index-based optical–SAR analysis, change maps).

1.2 Dev toolchain

Tool Version Notes
Python 3.14.3 (default) — also 3.12, 3.11 installed via py launcher Pin the project to 3.11 (py -3.11); PyTorch/rasterio/GDAL wheels lag on 3.14
pip 26.2 OK
uv 0.1.26 Present but old; upgrade or use plain pip + venv
Node v22.14.0 OK for Vite/React
npm 10.9.2 OK
Git 2.51.2.windows.1 OK
GitHub CLI 2.88.1 — logged in as KING-OF-FLAME, scopes gist, read:org, repo Auth ready
Docker not installed Not required for this build
Conda not on PATH Not required

Git identity: KING-OF-FLAME <mr.yashraj5233@gmail.com>

1.3 Repository state at start

  • No Git repository, no remote.
  • The only pre-existing entry was an empty satquery/ folder (zero files). It was present at the first scan and absent by the end of scaffolding; no delete command was issued against it. Nothing was lost — it held no content. The repo root is SIH/.

1.4 Colab compute backend (colab-proxy-mcp)

Connection opened and verified end-to-end: a cell was written into the live notebook, executed on the Colab runtime, and its stdout returned to the agent.

  • Round-trip proof: print("hello from Colab") produced hello from Colab
  • Runtime/GPU probe: see section 1.5.

The MCP connection is an agent-side authoring channel (it drives the notebook for us). It is not the app's runtime transport. The web app will talk to a FastAPI model server running inside the Colab runtime, exposed over a tunnel (cloudflared / ngrok) — cross-cutting infrastructure, see section 4.2.

1.5 Colab runtime probe

Executed live over the MCP channel on 2026-08-25:

Item Value
Python 3.13.15
OS Linux 6.6.122 x86_64, glibc 2.35
GPU NVIDIA Tesla T4 — 14.56 GB usable VRAM (15360 MiB), driver 580.82.07
Compute capability 7.5 (Turing)
PyTorch 2.11.0+cu128, cuda.is_available() == True
CPU / RAM 2 vCPU, 12.7 GB
Disk 113 GB total, 65 GB free

Consequences of this probe — these change the plan:

  1. VLM tier is confirmed at 3B, not 7B. Free-tier T4 with ~14.5 GB. Qwen2.5-VL-3B in 4-bit with LoRA fits with headroom; 7B would thrash.
  2. fp16, not bf16. Turing (cc 7.5) has no bf16 tensor cores. Training configs must use fp16=True with grad-scaling. A copied bf16 recipe will silently underperform or NaN.
  3. No FlashAttention-2 — it needs Ampere (cc 8.0+). Use PyTorch SDPA attention.
  4. Colab disk is 65 GB free, not 100 GB. The ~95 GB dataset budget in section 2.4 is a local F: drive budget and does not fit on the Colab runtime. Training must stream a stratified BigEarthNet subset from Drive or HF rather than materialising the full corpus on /content. This is an M1 design constraint, not an M3 surprise.
  5. 2 vCPU only — dataloader num_workers above 2 will hurt. Pre-tokenise and pre-resize into shards during M1 so the GPU is never waiting on CPU decode.
  6. Colab runs Python 3.13 while the local app is pinned to 3.11. Keep shared code in scripts/ free of version-specific syntax; the two environments only exchange JSON over HTTP, so the split is safe.

MCP channel note: update_cell and run_code_cell are reliable and fast. add_code_cell hung for ~10 minutes and had to be killed, which dropped the connection and reset the notebook. Workflow rule for the team: build notebooks by updating existing cells, not by inserting new ones.


2. Proposed tech stack

Optimised for: Colab GPU compute, 100 GB disk, a 24-hour build window, and the nine judging criteria.

2.1 Architecture in one line

React SPA → FastAPI (local, CPU) → Agentic controller → HTTP tunnel → FastAPI model server (Colab GPU) → specialist models

Split rationale: the judged deliverable is an interactive web app, which must stay responsive and demoable even if a Colab runtime dies. Keeping orchestration, validation, trace-building and report generation local — with only tensor work remote — means a runtime disconnect degrades the demo instead of killing it.

2.2 Component choices

Layer Choice Reasoning
Backend FastAPI + Uvicorn + Pydantic v2 Pydantic gives typed tool schemas and typed input-compatibility errors for free — exactly the "validate inputs / permitted parameters" requirement. WebSocket support streams the live execution trace, our strongest demo asset.
Async jobs FastAPI BackgroundTasks + in-process queue Celery/Redis is overkill for a 24 h build and a single-user demo.
Persistence SQLite + SQLModel Zero-ops run history and auditable execution traces. Judges can be shown past runs.
Geospatial I/O rasterio (GDAL), pyproj, shapely, numpy, OpenCV GeoTIFF/TIFF read, CRS handling, georeference and co-registration checks, windowed reads for large scenes.
DL runtime PyTorch 2.x + CUDA (Colab), HuggingFace transformers, PEFT (LoRA/QLoRA), TRL, bitsandbytes, accelerate LoRA on a T4 (16 GB) is the only way to fine-tune a VLM inside a hackathon window. PEFT adapters are ~100 MB, so they fit in Drive/HF and download fast.
RS domain adaptation CLIP ViT-B/32 contrastively fine-tuned on BigEarthNet.txt image–text pairs The mandatory adaptation requirement, satisfied with the prescribed dataset, trainable in 1–2 h on a T4. Low-risk and independently defensible in Q&A. Doubles as the task router's query encoder and as the confidence/similarity scorer.
VQA + captioning + grounding Qwen2.5-VL-3B-Instruct + LoRA on RSVQA / VRSBench One model covers three mandatory single-image capabilities; it emits bounding boxes natively, so text-guided grounding needs no second detector. 3B in 4-bit fits a T4 comfortably.
Change understanding TinyCD / ChangeFormer (bi-temporal siamese, LEVIR-CD pretrain) producing a change mask, then mask statistics narrated by the VLM; evaluated on CDVQA Splitting "where changed" (pixel model) from "what and how much changed" (VLM over mask statistics) is more accurate and more explainable than asking a VLM to eyeball two images. Also yields the optional spatial change map.
Optical–SAR fusion Physics-first: SAR despeckle (Lee filter) plus backscatter thresholding (water = low sigma-0, built-up = high sigma-0 / double-bounce) fused with optical NDWI/NDBI; a light learned refinement head on top Works without ISRO-specific labels, runs on CPU, and is trivially defensible to judges. Cartosat-2S / RISAT pairs at eval time are pre-georeferenced and co-registered, so pixel-aligned fusion is valid.
Agentic controller Custom typed controller — Registry → Planner → Validator → Executor → Synthesiser, emitting a JSON ExecutionTrace LangGraph/LangChain add dependency weight and hide the trace. Only the observable trace is scored, so a deterministic typed plan maximises score and debuggability. Routing = RS-CLIP text embeddings plus rules, with an optional LLM fallback below a confidence threshold.
Frontend React + Vite + TypeScript + TailwindCSS + shadcn/ui; MapLibre GL for georeferenced overlays; bi-temporal swipe slider UI/UX is criterion 7 and Prototype Quality is criterion 5 — a Streamlit/Gradio page reads as a mockup. Vite keeps the loop fast.
Debug UI Gradio, served from the Colab model server Free internal harness to test specialists before the React app is wired up. Never shown to judges.
Reports HTML template rendered to PDF via WeasyPrint, plus a JSON sidecar The "downloadable reports" deliverable; JSON keeps the trace machine-checkable.
Eval Custom harness in eval/, prescribed test splits, per-metric normalisation Mirrors the stated judging protocol so our numbers match theirs.
Testing pytest with tiny GeoTIFF fixtures Guards validators and routing — the parts most likely to break live on stage.
Quality ruff + black Fast, zero-config, no time tax.

2.3 Deliberate non-choices

  • No Docker — not installed, and containerising costs hours we do not have.
  • No Postgres / Redis / Celery — SQLite plus in-process background tasks suffice.
  • No LangChain / LangGraph — the trace is the product; hand-rolled beats framework-hidden.
  • No local training — 4 GB VRAM cannot hold a 3B VLM even in 4-bit with activations.
  • Python 3.11, not 3.14 — geospatial and DL wheel availability.

2.4 Disk budget (100 GB target, on F: with 797 GB free)

Item Estimate
BigEarthNet subset (patches plus BigEarthNet.txt captions) ~40 GB
RSVQA (LR plus an HR subset) ~10 GB
VRSBench ~12 GB
CDVQA plus LEVIR-CD ~8 GB
Base model weights plus LoRA adapters ~15 GB
Working intermediates, caches, reports ~10 GB
Total ~95 GB

This budget applies to the local F: drive only. The Colab runtime has just 65 GB free (section 1.5), so the full corpus cannot be staged there. Training reads a stratified BigEarthNet subset streamed from Drive/HF, pre-sharded during M1.


3. Risk register

Risk Impact Mitigation
Colab runtime disconnect or GPU quota exhaustion mid-demo High Cache every specialist's outputs for the demo images; ALLOW_LOCAL_FALLBACK path; pre-record a backup demo video
Tunnel URL rotates each Colab session Medium MODEL_SERVER_URL in .env plus a one-command re-point script
BigEarthNet download time Medium Start the fetch first (M1) and let it run in the background through M2–M3
Only 3.2 GB free RAM locally Medium Windowed rasterio reads; never load full scenes into memory
ISRO/SAC eval set unseen (Cartosat-2S plus RISAT) High Never hardcode to benchmark quirks; validators must accept arbitrary CRS, bit-depth and band-count GeoTIFFs
24 h window against five mandatory capabilities High Milestones are ordered so a working end-to-end vertical slice exists by M8; M9 is hardening and acceptance

4. Milestones

4.1 Unified scheme

This is the only numbering in use. Any older reference found in notes, commit messages or issues maps through 4.2.

ID Milestone PS requirement Status
M0 Scaffold, tool + trace contracts, dual backend PS 5 (contract) Done
M1 Dataset acquisition + geospatial I/O PS 6 Done — 19 GB, 36/36 checks
M2 Single-image VQA baseline PS 2 (VQA) Done — 5/5 coherent, 4/4 failure paths
M3 RS adaptation (LoRA) PS 1 (mandatory) Done — full run complete, section 13
M3-A Data ✅ Done — 11,763 records, section 8
M3-A2 Cross-modal records Done — 2,391 records appended, section 11
M3-B Smoke LoRA (tiny run, proves the loop) Done — section 12
M3-C Full LoRA adaptation Done — section 13 (dashboard 100 % / 25:27 elapsed; final adapter saved to Drive)
M3-D Eval Todo
M4 Captioning + grounding PS 2 (second task) Done — serving + tools implemented, unit-tested, and verified on 5 real VRSBench images — section 14
M5 Bi-temporal change analysis PS 3 (mandatory) Todo
M6 Optical–SAR cross-modal PS 4 (mandatory) Todo
M7 Agentic controller + execution trace PS 5 (mandatory) scaffolded (M0)
M8 Web UI + reports PS 7 scaffolded (M0)
M9 Hardening, docs, acceptance — Todo

Mandatory-capability coverage: RS adaptation → M3 · single-image VQA → M2 · second single-image task → M4 · change understanding → M5 · optical–SAR → M6 · agentic orchestration → M7 · GUI + reports → M8. All PS requirements are covered.

4.2 Old → new mapping

All older numbering maps to the unified scheme in one table. Commit messages are immutable, so they are interpreted through this table rather than rewritten.

Old A (original, M0–M11) Old B (re-scaffold, section 9.1) Previous canonical New unified Milestone
M0 Environment & scaffolding M0 Scaffold M0 M0 Scaffold, tool + trace contracts, dual backend
M1 Data acquisition M1 Datasets + geo I/O M1 M1 Dataset acquisition + geospatial I/O
M2 Colab GPU model-server harness M5 Colab FastAPI server cross-cutting cross-cutting Colab FastAPI server — see note below
M3 RS domain adaptation M3 RS adaptation M3 M3 RS adaptation (LoRA)
M4 Single-image VQA M2 VQA baseline M2 M2 Single-image VQA baseline
M5 Second single-image task — M4 M4 Captioning + grounding
M6 Bi-temporal change M6 Bi-temporal change M5 M5 Bi-temporal change analysis
M7 Optical–SAR analysis M7 Optical–SAR fusion M6 M6 Optical–SAR cross-modal
M8 Agentic orchestration M4 Agentic controller M7 M7 Agentic controller + execution trace
M9 Web application M8 Streamlit UI M8 M8 Web UI + reports
M10 Evaluation and reports M9 Benchmarks + QA M3-D / M9 M3-D / M9 Adapted-model eval → M3-D; final eval + reports → M9
M11 Hardening, demo, defence — M9 M9 Hardening, docs, acceptance

Two mappings need a word of explanation rather than a silent row:

  • The Colab FastAPI server is no longer its own milestone. It was M2 in scheme A and M5 in scheme B. It is now cross-cutting infrastructure: the contract (RemoteBackend, satquery/serving/api.py) was scaffolded in M0, and it gets completed when first genuinely needed — by M3-C for training-side serving and by M8 for the UI. It is not dropped; it simply is not a deliverable in its own right.
  • Old M10 / previous M3-D "Evaluation and reports" splits. Adapted-model evaluation belongs to M3-D (does the adapter actually help?). Benchmark runs over the prescribed test splits plus report generation belong to M9.

4.3 Already-committed work that used an old scheme

Commit messages are immutable, so they are mapped here rather than rewritten.

Commit Message says Actually is (unified)
c045a2f "M4/M5: Implement agentic controller, validators, classifier, confidence, colab server" M7 (controller + trace) plus the cross-cutting Colab server

When reading git history, treat any milestone number in a message dated 2026-08-25 or earlier as scheme A or B and resolve it through 4.2.

5. Open items

  • Create the GitHub remote (gh repo create) — auth is ready, the remote does not exist yet (awaiting a decision on repo name and visibility).
  • Confirm the Colab tier — free-tier Tesla T4 confirmed by probe. VLM tier fixed at 3B.
  • Decide whether to buy Colab Pro. A T4 will hold the plan, but an A100/L4 would cut M3 training time roughly 3-4x and remove the fp16-only constraint. Worth it if the 24 h window gets tight.
  • Confirm team size and the parallel work split for M4–M6 (they are independent).
  • Obtain BigEarthNet.txt — confirm the exact distribution/URL the problem statement refers to.

6. M1 — dataset acquisition (COMPLETE)

6.1 Where the data lives

Datasets live on the Colab runtime, not in this repo and not on the laptop. Rationale: all training and evaluation runs on the Colab GPU, so staging data next to the compute avoids a pointless 18 GB round trip over a home connection. The repo carries scripts and manifests only — .gitignore excludes data/, and this was verified with git check-ignore (section 6.6).

Path Contents
/content/data/raw/<job>/ Downloaded artefacts, one folder per job
/content/data/processed/ben_at_pairs/ Extracted co-registered S2+S1 patch pairs
/content/scripts/ Validation scripts (mirrors scripts/ in this repo)
/content/dl.log JSON-lines download log (per-file size and duration)

Colab runtimes are ephemeral and have reset twice already during M1. Re-acquisition is one command and takes ~75 s at observed speeds (50–100 MB/s):

python scripts/download_datasets.py --detach     # resumable; survives kernel restarts

6.2 What was acquired — 18.13 GB, 23 files, 0 failures

Job Source (HF) Size Subset decision
ben_txt BIFOLD-BigEarthNetv2-0/BigEarthNet.txt 0.45 GB Full annotation table — text only, so cheap to take whole
ben_s2_at yousunyu/BigEarthNet_S2_Austria 7.23 GB Austria only of ~60 GB full Sentinel-2
ben_s1_at seosiju/BigEarthNet-S1 5.07 GB Austria only of ~52 GB full Sentinel-1
vrsbench xiang709/VRSBench 3.83 GB Val images + the 3 official EVAL jsons; train images (7.97 GB) skipped
rsvqa_lr dmarsili/RSVQA-LR-2k 0.16 GB Pre-built 2k validation subset
rsvqa_hr dmarsili/RSVQA-HR-2k 0.90 GB Pre-built 2k validation subset
cdvqa ljx620/CDVQA 0.74 GB 10 of 1533 shards (1,000 samples) of a 114 GB corpus

Headroom after acquisition: 170 GB free on Colab, ~797 GB free on local F:. Model weights and checkpoints for M3–M6 need ~15 GB, so both budgets are comfortable.

6.3 The key find: BigEarthNet.txt is the backbone of this project

BigEarthNet.txt is the dataset the problem statement names, and it is far more useful than it first appears. It is 464,044 co-registered Sentinel-1 (SAR) + Sentinel-2 (multispectral) image pairs with 9,553,962 text annotations:

Annotation type Count Feeds milestone
binary 3,625,160 M2/M3 (VQA)
mcq 3,259,184 M2/M3 (VQA)
bounding box 2,205,686 M4 (text-guided grounding)
captioning 463,932 M4 (captioning)

Categories: presence, area, count, adjacency, point, reference, relative position, season, climate zone, country. It also carries latitude/longitude, so outputs can be geo-anchored.

One dataset therefore covers M3 (adaptation), M2 (VQA), M4 (captioning + grounding) and M6 (optical–SAR) — and it is the dataset the evaluators named. It also has a manually-verified bench split of exactly 1,082 patches / 15,029 annotations.

Because it is text-only, imagery had to be sourced separately, and the join key is the BigEarthNet v2.0 patch_id (S2) / s1_name (S1). Both joins were verified to work.

6.4 Subset strategy for BigEarthNet imagery

The full v2.0 imagery is ~110 GB and every HF mirror ships it as monolithic 5–46 GB archives with no selective access. Since the annotation table has a country column, Austria was chosen as a principled slice: 41,890 patches (train 23,817 / val 10,180 / test 7,853 / bench 40), available as country-scoped archives for both sensors at 12.3 GB total instead of 110 GB.

scripts/extract_bigearthnet_subset.py streams both archives once and keeps only patches present in both sensors and the annotation table, so every extracted pair is guaranteed usable for optical–SAR work.

Known limitation: the official bench split spans all countries, so evaluating on the full 1,082-patch benchmark would require the complete 110 GB download. Austria's test split (7,853 patches) is the working eval set; the 40 Austrian bench patches serve as a mini sanity benchmark. Revisit if Colab disk allows.

6.5 Validation results — 36/36 checks passed

One script per dataset, run on Colab (python scripts/validate_<name>.py):

Dataset Checks Samples Image dimensions Format
BigEarthNet 14/14 9,553,962 annotations / 464,044 patches S2: 8 bands (B02–B07, B11, B12); S1: VV + VH parquet + per-patch GeoTIFF
VRSBench 10/10 9,350 cap / 37,409 vqa / 16,159 referring 512×512 RGB, 9,350 val images json + zip
RSVQA 4/4 2,000 LR + 2,000 HR LR 256×256, HR 512×512 RGB parquet, images embedded
CDVQA 8/8 1,000 bi-temporal samples 512×512 RGB pairs webdataset tar

Format details worth carrying into later milestones:

  • VRSBench grounding boxes are encoded inside ground_truth as {<x1><y1><x2><y2>} with values normalised to 0–100 — not a separate box field. An obj_corner field carries pixel corners. All 16,159 referring records parsed. This convention is close to Qwen2.5-VL's own box format, which helps M4.
  • CDVQA has 8 question types (change_or_not 359, change_ratio_types 141, decrease_or_not 121, increase_or_not 113, smallest_change/largest_change 75 each, change_to_what 70, change_ratio 46) in a conversations schema with <image> placeholders — directly usable as VLM training/eval format. Source is CDVQA/SECOND.
  • RSVQA answers are heavily skewed to yes/no (LR: 690 no / 678 yes of 2,000), so accuracy alone will flatter a naive model. Report per-question-type accuracy in M9.
  • The S2 Austria tar was packed on macOS: ~43,800 AppleDouble ._* stubs and .DS_Store entries appear before any real file. Any extraction or dataloader must skip them or it will ingest garbage. Both validate_bigearthnet.py and extract_bigearthnet_subset.py filter them.

6.6 Git hygiene

data/ is excluded by .gitignore (patterns data/raw/*, data/interim/*, data/processed/*, plus *.tif, *.tar, *.zip, *.parquet, *.npy). Verified:

data/raw/foo.tif                 IGNORED
data/processed/shard.npy         IGNORED
models/checkpoints/a.safetensors IGNORED

No dataset file is tracked. Only scripts and manifests are committed.

6.7 Extracted optical–SAR pairs — verified co-registered

extract_bigearthnet_subset.py --n 1500 produced 1,500 complete pairs (458 MB) at /content/data/processed/ben_at_pairs/, each with 12 Sentinel-2 band files and 2 Sentinel-1 polarisations, plus pairs_manifest.json.

A pair opened with rasterio:

Sensor Files Grid dtype CRS Resolution
S2 B02/B03/B04/B08 120×120 uint16 EPSG:32633 10 m
S2 B05/B06/B07/B8A/B11/B12 60×60 uint16 EPSG:32633 20 m
S2 B01/B09 20×20 uint16 EPSG:32633 60 m
S1 VV, VH 120×120 float32 EPSG:32633 10 m

Sample values: S2 B04 min=120 max=2702 mean=524 (reflectance DN); S1 VV min=-20.36 max=11.59 mean=-8.73 (dB backscatter — physically sensible).

Both sensors share EPSG:32633 at 10 m on a 120×120 grid, so the optical and SAR rasters are pixel-aligned with no reprojection needed. This is exactly the co-registered optical–SAR configuration M6 requires, and it means the fusion work can start from real data rather than synthetic alignment. The 20 m and 60 m S2 bands need upsampling to the 10 m grid before fusion.


7. COMPUTE STATUS

Detected 2026-08-25 on the connected Colab Pro runtime via python training/compute_check.py --benchmark --write --apply. All values measured, none assumed.

Item Status
Connected PASS — Colab MCP executes remotely end-to-end
GPU PASS — NVIDIA A100-SXM4-80GB, 1 device, compute capability 8.0
VRAM 79.25 GB (81920 MiB reported by nvidia-smi)
CUDA 12.8 runtime, driver 580.82.07
PyTorch 2.11.0+cu128 — CUDA build, cuda.is_available() == True
CPU / RAM 12 vCPU / 167.05 GB
Disk 183 GB free of 235.7 GB
Drive FAIL — not mounted (only remaining blocker)
Mixed precision ENABLED — bf16 (cc 8.0 has bf16 tensor cores)
TF32 ENABLED — applied to matmul + cuDNN
Benchmark 243.3 TFLOPS bf16 (8192² matmul × 32); peak alloc 0.38 GB
Ready for LoRA training YES — --require-gpu exits 0

Recommended config (auto-derived from measured VRAM)

device: cuda            dtype: bfloat16         mixed_precision: bf16
batch_size: 8           gradient_accumulation_steps: 2    effective_batch_size: 16
gradient_checkpointing: false     num_workers: 4    pin_memory: true
load_in_4bit: false     tf32: true    attn_implementation: sdpa
max_visual_tokens: 1280           distributed_strategy: single

79.25 GB lands in the top VRAM tier, so gradient checkpointing and 4-bit quantisation are both off — neither is needed, and both cost speed or quality. Effective batch is held at 16 to match the trajectory used on smaller tiers. Single GPU, so no distributed strategy.

Resolved since the previous check

Both earlier blockers are cleared. The runtime was CPU High-RAM with a CPU-only torch wheel; it is now an A100 with +cu128. Nothing in the code changed to achieve this — the detector simply reported the new hardware, which is the point of not hard-coding a GPU.

One correctness fix this check surfaced

The recommender initially proposed flash_attention_2 purely from compute capability 8.0. flash_attn is not installed on this runtime, so that config would have crashed model loading. Hardware capability is necessary but not sufficient; the recommendation now requires the silicon and the import, and falls back to sdpa. The backend re-checks at load time as a second guard. Installing flash-attn --no-build-isolation would unlock it and give a further speed-up; sdpa is correct and safe meanwhile.

Remaining blocker: Drive not mounted

/content is ephemeral — this project has already lost it four times in one session. LoRA adapters and checkpoints from M3-C are the irreplaceable artefacts; losing them costs GPU-hours. Mount before any long run (interactive OAuth, so a human must run it):

from google.colab import drive
drive.mount('/content/drive')
import training.drive_setup as ds
print(ds.setup(mount=False))     # creates MyDrive/SatQueryAI/{datasets,checkpoints,adapters,results,logs,cache}

This does not block M3-A (data preparation, re-derivable). It does block M3-C.

Guardrail status

--require-gpu exits 0 on this runtime (training permitted) and exited 1 on both the CPU Colab runtime and the local Windows machine. The guard is verified in both directions, not just the failing one.


8. M3-A — BigEarthNet image-text adaptation data (COMPLETE)

Built and validated 2026-08-25 on the A100 runtime. No training was run. The M2 baseline was not modified.

8.1 What was built

data/processed/bigearthnet_image_text.jsonl — 11,763 instruction records over 1,000 image patches. This is instruction supervision, not classification: every record is {sample_id, image_path, question, answer, source, split} plus patch_id, labels, question_kind, supervision and sar_image_path for auditability.

8.2 Every answer comes from real data

Two authoritative sources; no label was invented.

Source Records What is real
metadata.parquet (BigEarthNet v2.0, official) 7,000 CORINE multi-labels, official split, country, snow/cloud flags. Question wording is templated; every answer is read off the actual fields
BigEarthNet.txt.parquet (BIFOLD) 4,763 Question and answer both authored by the dataset creators — captioning, binary, multiple-choice

The 14-class label vocabulary is derived from the data itself, not hard-coded.

8.3 Splits and leakage

Splits are the official BigEarthNet v2.0 splits, taken from metadata.parquet. Split is a property of the patch, so every question generated from a patch inherits that patch's split — a patch cannot span splits by construction, and validate_manifest.py re-checks it independently against both patch_id and image_path.

Split Patches Records
train 600 7,025
validation 200 2,365
test 200 2,373

8.4 Determinism

Patch selection is sorted(patch_id)[:N] per official split; every per-patch random choice is seeded from sha256(patch_id).

Proven twice, not asserted:

  1. Same machine — rebuilt on Colab, byte-identical file SHA-256 (b697f6c15dd4382f...).
  2. Across machines and Python versions — independently regenerated on the local Windows mirror (Python 3.14.3) from the local dataset copy, and the content hash matched Colab (Linux, Python 3.13.15) exactly: 897235331c3eae3cf6c5191e17e6daa08f0bea192214cd858ac6dade5d377006 (over sample_id|question|answer|split|question_kind; absolute image paths differ by machine and are excluded). The local validator also exits 0 with all checks green.

8.5 Validation — all checks green (exit 0)

[PASS] schema: all 6 required fields present
[PASS] splits: no patch spans splits (1000 patches)
[PASS] splits: no image path spans splits
[PASS] ids: 11,763 unique sample_id
[PASS] no duplicate (patch, question) pairs
[PASS] questions: non-empty, <= 2000 chars
[PASS] answers: non-empty, <= 4000 chars
[PASS] answers agree with real labels (4,000 cross-checked)
[PASS] negative-presence questions name absent classes
[PASS] images: 1,000 unique dirs exist with RGB bands
[INFO] paired SAR dirs present: 1,000/1,000

The validator re-derives its checks from the manifest and source metadata rather than trusting the builder, so a builder bug is caught independently.

8.6 Distributions

Question kinds: bentxt_mcq 1,906 · bentxt_binary 1,904 · landcover_list 1,000 · landcover_count 1,000 · presence_positive 1,000 · presence_negative 1,000 · snow_flag 1,000 · cloud_flag 1,000 · country 1,000 · bentxt_captioning 953

Labels (per patch, 14 classes): Pastures 768 · Mixed forest 615 · Urban fabric 455 · Agriculture-with-natural-vegetation 323 · Coniferous forest 300 · Arable land 293 · Broad-leaved forest 283 · Complex cultivation 245 · Inland waters 213 · Industrial/commercial 68 · Inland wetlands 35 · Natural grassland 22 · Transitional woodland 19 · Moors/heathland 8

8.7 Two characteristics that must shape M3-C and M3-D

  1. Answer polarity is skewed 2.08:1 — no 3,985 vs yes 1,919. This comes from snow_flag/cloud_flag being overwhelmingly "no" on a cloud-filtered Austrian subset, on top of the deliberate 1:1 presence pairing. Left uncorrected a model can score well by answering "no". M3-C should weight or rebalance binary answers, and M3-D must report per-question-kind accuracy — aggregate accuracy will flatter a degenerate model.
  2. Labels are long-tailed, 96:1 (Pastures 768 → Moors 8). This is real Austrian land cover and was deliberately not resampled; rare-class performance must be read per class.

Also noted: 47 of 1,000 patches carry no BigEarthNet.txt annotation, because that table covers 464,044 of the 480,038 patches in metadata. Those patches still contribute their 7 metadata-grounded questions.

8.8 Artefacts

Path Committed?
training/geochat_adaptation/build_manifest.py yes
training/geochat_adaptation/validate_manifest.py yes
training/geochat_adaptation/dataset_report.json yes
training/geochat_adaptation/sync_to_drive.py yes
scripts/extract_ben_patches.py yes
data/processed/bigearthnet_image_text.jsonl no — git-ignored data
extracted patches (~306 MB) no — git-ignored data

Regenerate anywhere in ~4 minutes:

python scripts/extract_ben_patches.py --raw data/raw --out data/processed/ben_patches
python training/geochat_adaptation/build_manifest.py --raw data/raw --patches data/processed/ben_patches
python training/geochat_adaptation/validate_manifest.py

8.9 Drive status — still NOT mounted

Reported rather than assumed: /content/drive/MyDrive does not exist on this runtime, so MyDrive/SatQueryAI/ could not be created or written. The manifest currently lives on VM disk and on the local mirror. sync_to_drive.py is ready and will copy the manifest, selection and report (with SHA-256s) as soon as Drive is mounted:

from google.colab import drive; drive.mount('/content/drive')
import training.drive_setup as ds; ds.setup(mount=False)
!python training/geochat_adaptation/sync_to_drive.py

This does not block M3-B (LoRA smoke test), whose inputs are re-derivable. It does matter before M3-C, whose adapter output is not.


9. Re-scaffold: satquery/ package, dual backend, trace contract

Added 2026-08-25 on top of the existing M0–M3-A work. Additive: nothing that worked was deleted or overwritten.

9.1 Milestone numbering

Superseded. The unified milestone scheme now lives in section 4.1, with the single historical mapping in section 4.2. Nothing here is authoritative any more.

9.2 The ExecutionTrace contract

satquery/core/schemas.py. PS 5 says only the observable trace is evaluated, so this is treated as the most important code in the repo. 21 top-level fields, every PS-named item first-class rather than a free-form blob:

trace_id · timestamp · schema_version · query · input_paths · classified_task · task_confidence · input_config · classifier_method · input_validation · plan · steps · final_answer · visual_evidence_paths · overall_confidence · backend · rs_adapted · total_duration_ms · status · warnings · errors

Verified, not assumed: round-trips through model_dump_json() → model_validate_json() byte-identically; confidences outside [0,1] are rejected by validators.

Three deliberate choices:

  • task_confidence and overall_confidence are separate. "Confident this is a change query, not confident in the change estimate" is the honest message.
  • rs_adapted is on the trace itself. PS 1 fails a generic VLM, so a run must state whether it was adapted rather than leaving it implicit.
  • A failed run is still a valid trace (status="failed"), never a lost run.

9.3 Dual backend — the architecture constraint

satquery/core/backends.py: LocalBackend | RemoteBackend, selected by MODEL_BACKEND via get_backend(). Remote is an HTTP client for the Colab FastAPI server behind cloudflared; the tunnel URL is config (SATQUERY_REMOTE_URL) because it rotates every session.

Enforced by tests/test_backend_abstraction.py: any import transformers/torch/peft outside the six permitted modules fails the suite. A stray import would work locally and silently stop using the A100 in remote mode — an error-free bug, the worst kind. The package also imports cleanly with no ML stack installed.

9.4 Collisions with existing code (deliberately unresolved)

satquery/agent/ vs top-level agent/, and satquery/training/ vs training/. Both are live. The old paths hold working, tested code (M2 VQA, M3-A pipeline, compute detection, 26 tests); the new ones are stubs. Migration lands with M4, the first milestone that adds a new tool and therefore has to import both — a rewrite now would throw away validated work to satisfy a directory layout.

Same for .env.example: the new MODEL_BACKEND / SATQUERY_REMOTE_URL / ADAPTER_PATH vars were appended, leaving the existing keys intact.


10. Blocked / Risks

Updated every milestone.

# Risk Severity State
1 F: is exFAT on a USB SSD, volume health Warning critical RESOLVED — repo re-cloned to D:\SatQueryAI (NTFS internal drive)
2 Google Drive not mounted on Colab high ACTIVE
3 Colab runtime resets wipe /content medium mitigated
4 RS adaptation (PS 1) not yet trained high in progress
5 flash_attn not installed on Colab low mitigated
6 Answer polarity skewed 2.08:1 toward "no" medium resolved by M3-A2 (now 1.37:1)
7 Milestone numbering drift low unified (4.1 / 4.2)

Risk 1 — data loss on F: (happened, recovered)

At 23:24 the entire working tree vanished: all 63 tracked files plus 19 GB of datasets. No delete command was issued. Earlier the same session, an empty satquery/ folder and the whole .venv disappeared the same way.

Diagnosis:

F:  exFAT   USB   Seagate One Touch SSD   HealthStatus: Warning
C:  NTFS    NVMe                          Healthy
D:  NTFS    SATA                          Healthy   288.6 GB free
E:  NTFS    SATA                          Healthy   305.3 GB free

exFAT has no journaling. A momentary USB disconnect corrupts the directory table and Windows drops whole trees. Not antivirus — Defender's newest detection is from 2024. This also explains git's "dubious ownership" warning in section 1.2 (exFAT records no ownership).

Recovered fully: .git survived, HEAD == origin/main, git checkout -- . restored 63/63 files, git fsck clean, 26/26 tests pass. Only the 19 GB of datasets was truly lost, and that re-downloads in ~75 s on Colab.

Mitigation — do this before more work: move the project to D: (NTFS, internal SATA, 288 GB free):

git clone https://github.com/KING-OF-FLAME/satquery-ai.git D:/satquery-ai

Until then: commit and push after every meaningful step, and treat local data as disposable. GitHub is the only durable copy.


11. M3-A2 — cross-modal and rebalanced supervision

Real-data run completed and validated 2026-08-26 on the connected A100-80GB Colab runtime. M3-A records are untouched; M2 is untouched.

11.1 What it adds (second pass, appended)

training/geochat_adaptation/crossmodal.py plus a second pass in build_manifest.py. The first pass is never rebuilt.

Kind Target Real count Why
cross_modal ~1,000 1,000 PS 4. Not one M3-A record showed both sensors at once, yet every patch has a verified co-registered S1 partner (1,000/1,000). Two images per record: S2 optical + S1 SAR.
sar_only ~400 391 With optical in every record a model can ignore the SAR branch entirely and still score well. SAR-only records make that shortcut unprofitable.
presence_positive_rebalance computed 1,000 Corrects the 2.08:1 "no" skew M3-A measured. Positives are added; nothing is deleted.

The rebalance count is computed from the measured first-pass polarity (need = no/target - yes), not hard-coded. The measured deficit was 1,146, but at most one positive can be added per patch, so 1,000 positives were added and the ratio moved from 2.08:1 to 1.37:1.

11.2 No fact is invented

Question wording and the sensor-attribution phrasing are templated; every claim about scene content is gated on the patch's real CORINE label array. Verified across four real label combinations — built-up+water, neither, water-only, built-up-only — that the answer asserts built-up/water iff the labels contain it, and that no class name appears in an answer unless it is in that patch's label array.

The attribution describes how each sensor evidences a class the labels say is genuinely present (SAR double-bounce for structure, specular return for water, optical NIR for vegetation). It never asserts a class the labels lack.

11.3 Schema extension

images is the canonical multi-image field (1 or 2 paths). image_path remains populated on every record — M2 and the existing loader read it and must not break. The validator enforces images[0] == image_path, so the two can never drift apart silently.

11.4 Validation results — all checks green (exit 0)

[PASS] schema: all 6 required fields present
[PASS] splits: no patch spans splits (1000 patches)
[PASS] splits: no image path spans splits
[PASS] ids: 14,154 unique sample_id
[PASS] no duplicate (patch, question) pairs
[PASS] questions: non-empty, <= 2000 chars
[PASS] answers: non-empty, <= 4000 chars
[PASS] answers agree with real labels (4,000 cross-checked)
[PASS] negative-presence questions name absent classes
[PASS] every record has `images` (1 or 2 paths)
[PASS] images[0] == image_path (backward compatible)
[PASS] cross_modal: exactly 2 images (1000 records)
[PASS] cross_modal: one S2 + one S1, same patch_id
[PASS] cross_modal answers agree with real labels (1,000 checked)
[PASS] sar_only: SAR image only (391 records)
[PASS] images: 2,000 unique dirs exist with RGB bands
[INFO] paired SAR dirs present: 1,000/1,000
VALIDATION PASSED: all checks green

11.5 Counts and distributions

Metric Value
Total records 14,154
M3-A first pass 11,763
M3-A2 second pass 2,391
Patches 1,000
cross_modal 1,000
sar_only 391
presence_positive_rebalance 1,000
Answer polarity no=3,985 · yes=2,919 · 1.37:1 (was 2.08:1)
Records per split train 8,472 · validation 2,838 · test 2,844
Patches per split train 600 · validation 200 · test 200

11.6 Determinism re-proven

Two consecutive Colab builds produced an identical content hash: e5656e0d640b488291564443c259d1ee333d95534b1b653fcdfd846d77aa31ab over sample_id|question|answer|split|question_kind. Absolute image paths are excluded from the hash because they differ by machine.

11.7 Two full cross_modal sample records

Record 1 — no built-up or water, but vegetation and agriculture present:

{
  "sample_id": "ben-S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_26_57-xmodal-00",
  "images": [
    "/content/data/processed/ben_patches/s2/S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_26_57",
    "/content/data/processed/ben_patches/s1/S1B_IW_GRDH_1SDV_20170612T165809_33UUP_26_57"
  ],
  "image_path": "/content/data/processed/ben_patches/s2/S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_26_57",
  "sar_image_path": "/content/data/processed/ben_patches/s1/S1B_IW_GRDH_1SDV_20170612T165809_33UUP_26_57",
  "modalities": ["optical", "sar"],
  "question": "Combine the optical and SAR observations to describe what this area contains.",
  "answer": "No built-up class is recorded for this scene, and the SAR image shows no concentrated high-backscatter structure to indicate one. No water class is recorded, and no extended low-backscatter region consistent with open water is indicated. Vegetated cover is present (Broad-leaved forest, Mixed forest), which the optical bands separate by their high near-infrared reflectance while SAR contributes volume scattering from the canopy. Agricultural cover is present (Arable land, Pastures); field structure is visible optically and SAR adds surface-roughness contrast between parcels. The two sensors are complementary here: SAR carries structure and is unaffected by cloud, while the optical bands carry the spectral signature needed to separate land-cover types.",
  "question_kind": "cross_modal",
  "supervision": "metadata_template_crossmodal",
  "source": "BigEarthNet",
  "split": "test",
  "patch_id": "S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_26_57",
  "s1_name": "S1B_IW_GRDH_1SDV_20170612T165809_33UUP_26_57",
  "labels": ["Arable land", "Broad-leaved forest", "Mixed forest", "Pastures"],
  "country": "Austria"
}

Record 2 — water present but no built-up:

{
  "sample_id": "ben-S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_27_55-xmodal-00",
  "images": [
    "/content/data/processed/ben_patches/s2/S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_27_55",
    "/content/data/processed/ben_patches/s1/S1B_IW_GRDH_1SDV_20170612T165809_33UUP_27_55"
  ],
  "image_path": "/content/data/processed/ben_patches/s2/S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_27_55",
  "sar_image_path": "/content/data/processed/ben_patches/s1/S1B_IW_GRDH_1SDV_20170612T165809_33UUP_27_55",
  "modalities": ["optical", "sar"],
  "question": "Combine the optical and SAR observations to describe what this area contains.",
  "answer": "No built-up class is recorded for this scene, and the SAR image shows no concentrated high-backscatter structure to indicate one. Water is present (Inland waters). Open water is a specular reflector, so it returns very low SAR backscatter and appears dark, while the optical bands show the characteristic strong absorption in the near-infrared. Vegetated cover is present (Broad-leaved forest, Mixed forest), which the optical bands separate by their high near-infrared reflectance while SAR contributes volume scattering from the canopy. Agricultural cover is present (Arable land); field structure is visible optically and SAR adds surface-roughness contrast between parcels. The two sensors are complementary here: SAR carries structure and is unaffected by cloud, while the optical bands carry the spectral signature needed to separate land-cover types.",
  "question_kind": "cross_modal",
  "supervision": "metadata_template_crossmodal",
  "source": "BigEarthNet",
  "split": "test",
  "patch_id": "S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_27_55",
  "s1_name": "S1B_IW_GRDH_1SDV_20170612T165809_33UUP_27_55",
  "labels": ["Arable land", "Broad-leaved forest", "Inland waters", "Mixed forest"],
  "country": "Austria"
}

12. M3-B — LoRA smoke run

Completed and verified 2026-08-26 on the connected A100-80GB Colab runtime. M3-A/A2 records and M2 are untouched. This was an insurance run to prove the full training → save → reload → inference pipeline before committing hours of GPU time to M3-C.

12.1 What was built

  • training/geochat_adaptation/train_lora.py — config-driven LoRA trainer.
  • training/geochat_adaptation/configs/m3b_smoke.yaml — smoke-run hyperparameters.
  • training/geochat_adaptation/verify_adapter.py — base + adapter inference test.

12.2 Compute routing

The trainer consumes config/compute.yaml via training.compute_config:

Compute: NVIDIA A100-SXM4-80GB | 79.25 GB VRAM | bfloat16 | bs=8x2=16 | workers=4
Source: config
Applied: tf32 enabled (matmul + cudnn)
Applied: cudnn.benchmark enabled
Applied: float32_matmul_precision=high

12.3 LoRA configuration

  • Base: Qwen/Qwen2.5-VL-3B-Instruct
  • LoRA r=16, alpha=32, dropout=0.05
  • Targets: q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
  • Vision tower frozen
  • Trainable params: 37,152,768 / 3,791,775,744 total (0.9798%)

The original config also targeted the vision-language merger, but PEFT raised Target module ... is not supported because Qwen2.5-VL's patch merger uses custom module types. The language tower LoRA surface is the supported, verifiable target.

12.4 Multi-image collator verification

The sanity pass explicitly exercises single-image, cross_modal (2-image), and sar_only records before training starts:

Sanity pass: one single-image record...
  single-image loss=2.2153  input_ids=(1, 68)  supervised_tokens=21
Sanity pass: one cross_modal (2-image) record...
  cross_modal loss=2.9684  input_ids=(1, 238)  images_in_record=2  supervised_tokens=167
Sanity pass: one sar_only record...
  sar_only loss=3.2347  input_ids=(1, 109)  supervised_tokens=58

The cross_modal input is significantly longer because the processor receives both the optical and SAR images, confirming the collator does not silently drop the second image.

12.5 Training log tail (Drive)

Adapter written to /content/drive/MyDrive/SatQueryAI/adapters/m3b-smoke-6d6823eb95e6/.

{"step": 10,  "epoch": 0.317, "loss": 1.885, "grad_norm": 1.197, "learning_rate": 8.445e-05}
{"step": 20,  "epoch": 0.635, "loss": 0.973, "grad_norm": 1.033, "learning_rate": 3.747e-05}
{"step": 30,  "epoch": 0.952, "loss": 0.885, "grad_norm": 1.567, "learning_rate": 2.293e-06}
{"step": 32,  "epoch": 1.0,   "train_loss": 1.243, "train_runtime": 41.0s}

Loss is not the point of a smoke run; the point is that the pipeline runs, back-propagates, saves, and reloads.

12.6 Adapter file listing on Drive

total 153M
-rw------- 1 root root 1.2K  adapter_config.json
-rw------- 1 root root 142M  adapter_model.safetensors
-rw------- 1 root root 1.3K  processor_config.json
-rw------- 1 root root  11M  tokenizer.json
-rw------- 1 root root 5.1K  README.md
-rw------- 1 root root  967  training_config.json
-rw------- 1 root root 899   training_log.jsonl
drwx------ 2 root root 4.0K  checkpoint-32

12.7 verify_adapter.py outputs

Single-image record:

{
  "sample_id": "ben-S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_26_57-meta-00",
  "question_kind": "landcover_list",
  "question": "What land-cover categories are present in this scene?",
  "expected_answer": "Arable land, Broad-leaved forest, Mixed forest, Pastures",
  "generated_answer": "arable land, forest, pasture",
  "n_images": 1
}

Cross_modal (2-image) record:

{
  "sample_id": "ben-S2A_MSIL2A_20170613T101031_N9999_R022_T33UUP_26_57-xmodal-00",
  "question_kind": "cross_modal",
  "question": "Combine the optical and SAR observations to describe what this area contains.",
  "expected_answer": "No built-up class is recorded ...",
  "generated_answer": "The area is dominated by agricultural land use, with some forest areas.",
  "n_images": 2
}

Both single-image and 2-image records produce coherent text. The 2-image record specifically proves the SAR branch reaches the model; a single-image collator would have failed here.

12.8 Fixes applied during M3-B

  1. Dropped unsupported merger target from LoRA config.
  2. Added warmup_steps fallback for older TrainingArguments.
  3. Made LoggingCallback inherit from TrainerCallback.
  4. Ensured the sanity pass searches the full manifest for a real cross_modal record, so --max_samples caps do not hide multi-image examples.

13. M3-C — full RS adaptation runs

Two full runs were launched under this milestone, and both are recorded below rather than one being deleted, because the adapter that ships is a descendant of the second and the provenance chain has to be auditable:

Run name Role Kept artefacts
m3c-full-6d20f38b53eb first full launch (§13.A) checkpoints 600-1060
m3c-full-a6796d97783f0fec the run that produced the shipped adapter checkpoints 1060-1590

The adapter used everywhere else in this project is m3c-full-a6796d97783f0fec/checkpoint-1590 (see §13.6 for the M3-C2 resume that produced it and §14 for its evaluation). Sections 13.1-13.4 describe that run's configuration; §13.A holds the first launch's record; §13.5 onward covers results, the resume, and the overfitting finding.

Launched 2026-08-26T07:49:26Z on the connected A100-80GB Colab runtime via nohup, so it survives session breaks.

13.1 Run configuration

Setting Value
Manifest data/processed/bigearthnet_image_text.jsonl (14,154 records)
Train split 8,472 records (60%)
Validation split 2,838 records (20%)
Epochs 2
Effective batch 8 × 2 = 16
Estimated steps 1,060
Wall-clock estimate ≈27 minutes (0.45 h), well under the 2.5 h deadline
LR 1e-4 cosine with 3% warmup
Save interval every 200 steps to Drive
Eval interval every 500 steps on validation split
Run name m3c-full-a6796d97783f0fec
Adapter output /content/drive/MyDrive/SatQueryAI/adapters/m3c-full-a6796d97783f0fec
Log /content/drive/MyDrive/SatQueryAI/logs/m3c-full-a6796d97783f0fec_train.log

13.2 Guardrails

  • Checkpoints save to Drive only, never to /content.
  • NaN loss detection: if any reported loss is NaN, the trainer aborts and writes NAN_ABORT marker next to the adapter.
  • Provenance: the manifest SHA-256 (a6796d97783f0fec) is recorded in manifest_sha256.txt inside the run directory.
  • Vision tower remains frozen; only language attention + MLP LoRA is trainable.

13.3 Live monitoring commands

# (a) watch progress
tail -f /content/drive/MyDrive/SatQueryAI/logs/m3c-full-a6796d97783f0fec_train.log

# (b) kill if needed
kill $(cat /content/drive/MyDrive/SatQueryAI/logs/m3c-full-a6796d97783f0fec.pid)

# (c) adapter path
ls -lh /content/drive/MyDrive/SatQueryAI/adapters/m3c-full-a6796d97783f0fec

13.4 Launch command

cd /content/satquery-ai
nohup bash /content/drive/MyDrive/SatQueryAI/logs/m3c-full-a6796d97783f0fec_run.sh   >/dev/null 2>&1 &
echo $! > /content/drive/MyDrive/SatQueryAI/logs/m3c-full-a6796d97783f0fec.pid

13.A First full launch (m3c-full-6d20f38b53eb)

Launched 2026-08-26 on the connected A100-80GB Colab runtime as a background process so it survives this session. Retained for provenance; the shipped adapter comes from the run in §13.1-13.4.

13.A.1 Wall-clock estimate (before launch)

Item Value
Training records 8,472
Batch size 8
Gradient accumulation 2
Effective batch size 16
Steps per epoch 530
Epochs 2
Total steps 1,060
Observed step time (M3-B) ~1.3 s/step
Estimated wall-clock ~23 minutes

The 2-epoch run fits well inside the 2.5-hour hard deadline, so it was launched at full length rather than reduced to 1 epoch.

13.A.2 Launch details

python training/geochat_adaptation/train_lora.py \
  --config training/geochat_adaptation/configs/m3c_full.yaml \
  --manifest /content/data/processed/bigearthnet_image_text.jsonl \
  --require-gpu \
  --time-limit-minutes 150
  • PID: 40019
  • Run name: m3c-full-6d20f38b53eb
  • Adapter output path: /content/drive/MyDrive/SatQueryAI/adapters/m3c-full-6d20f38b53eb
  • Progress log: tail -f /content/drive/MyDrive/SatQueryAI/logs/m3c-full-6d20f38b53eb.out
  • JSON training log: tail -f /content/drive/MyDrive/SatQueryAI/adapters/m3c-full-6d20f38b53eb/training_log.jsonl
  • Kill command: kill -TERM 40019

13.A.3 Guardrails active

  1. Drive checkpoints only — save_steps: 200, save_total_limit: 4; output directory is under /content/drive/MyDrive/SatQueryAI/adapters/.
  2. NaN abort — a guardrail callback stops training and writes {output_dir}/abort_marker.txt if loss becomes NaN.
  3. Wall-clock hard cap — --time-limit-minutes 150 stops the run after 2.5 hours and writes the marker file.
  4. Manifest provenance — manifest_sha256.txt is written next to the adapter so the exact training data can be reproduced/audited.

13.A.4 What to watch for

  • training_log.jsonl should show loss falling smoothly (M3-B smoke went 1.885 → 0.973 → 0.885 over one epoch).
  • First eval on the validation split will land around step 500.
  • Checkpoints will be written at steps 200, 400, 600, 800, 1000.
  • If abort_marker.txt appears, read its reason line.
  • The run should finish in ~23 minutes.

13.5 Results

Completed 2026-08-26T06:11Z on A100-80GB Colab.

Metric Value
Total steps 1,060 / 1,060
Wall-clock 25 min 46 s
Train loss 0.3019
Eval loss 0.3021
Eval samples/sec 33.24
Final adapter adapter_model.safetensors (141.82 MB)
Checkpoints saved 200, 400, 600, 800, 1000, 1060

verify_adapter.py confirms base + adapter loads and answers both single-image and cross_modal (2-image) records successfully.

13.6 M3-C2 — resume from checkpoint-1060 for one additional epoch

Resumed 2026-08-26 on the connected A100-80GB Colab runtime from checkpoint-1060 (global step 1060, epoch 2.0). Because Hugging Face Trainer resumes steps/epochs from trainer_state.json, the run was launched with --epochs 3 so it completed exactly one additional epoch.

Metric Value
Resumed from checkpoint-1060
Final checkpoint checkpoint-1590
Total steps after resume 1,590
Wall-clock for the extra epoch 13 min 14 s
Train loss 0.06044
Eval loss 0.3057
New checkpoints saved 1200, 1400, 1590

verify_adapter.py (5 records — the previously failing single-image and cross-modal records plus 3 deterministic random records) result:

Check Result
Pure non-CORINE hallucination 0 / 5 (fixed)
CORINE-class over-listing 2 / 5
Model answers non-landcover questions correctly 3 / 3 (cloud flag, MCQ, presence)

Side-by-side for the two originally failing records:

  • Single-image landcover_list — true labels: Arable land, Broad-leaved forest, Mixed forest, Pastures. Generated: Arable land, Complex cultivation patterns, Coniferous forest, Pastures, Urban fabric. Over-listed (in vocab, not true): Complex cultivation patterns, Coniferous forest, Urban fabric.
  • Cross-modal cross_modal — true labels: Arable land, Broad-leaved forest, Mixed forest, Pastures. Generated answer is now a well-structured sensor-fusion paragraph, but still over-lists Complex cultivation patterns, Urban fabric.

Honest verdict: the extra epoch eliminated pure hallucination (no made-up class names), which was the M3-C2 stop condition. The model still occasionally cites extra CORINE classes that are not in a patch's true label set, so factual precision on class-list tasks is not perfect. Per the stop rule, no epoch 4 was launched automatically.

The final adapter remains at: /content/drive/MyDrive/SatQueryAI/adapters/m3c-full-a6796d97783f0fec/adapter_model.safetensors (141.8 MB).

13.7 Reproducing the adapter locally

The adapter is git-ignored because it is a large binary artefact. To copy it into the repo for local runs:

# On the Colab runtime where the adapter was produced:
bash scripts/download_m3c2_adapter.sh

This writes to models/adapters/m3c-full-a6796d97783f0fec/. That directory is ignored, but the script is committed and the Drive path is durable, so the artefact is always reproducible.

13.8 Honest overfitting finding

The M3-C2 extra epoch is accepted as the final adapter. The numbers make the reason explicit:

Metric M3-C (2 epochs) M3-C2 (3 epochs) Interpretation
Train loss 0.3019 0.06044 Model memorised the training split
Eval loss 0.3021 0.3057 Held-out performance did not improve; it slightly worsened

This confirms the stop rule fired correctly. The extra epoch removed pure (non-CORINE) hallucination, but it did not remove CORINE-class over-listing, and further epochs on this data would almost certainly increase overfitting without fixing the residual precision error. Any future improvement needs a data or objective change, not more epochs.

14. M3-D — adapter evaluation and adaptation evidence

Scope note: M3-D does not retrain. checkpoint-1590 is accepted as the final adapter; this milestone measures it and documents the result honestly.

14.1 The final adapter

Item Value
Run m3c-full-a6796d97783f0fec (M3-C2, resumed from checkpoint-1060)
Accepted checkpoint checkpoint-1590 — epoch 3.0
Drive location /content/drive/MyDrive/SatQueryAI/adapters/m3c-full-a6796d97783f0fec
Local location models/adapters/m3c-full-a6796d97783f0fec (git-ignored)
Trainable parameters 37,152,768 across 696 LoRA tensors, fp32
Adapter file adapter_model.safetensors, 148,712,776 bytes
SHA-256 3b625d63ed4f28640a6fb14072486305…
Training manifest SHA-256 a6796d97783f0fecf091688dddf41c728454c964121e64b1e9a60fb0da703d8f

Verified locally: the top-level adapter_model.safetensors is byte-identical to checkpoint-1590/adapter_model.safetensors (same SHA-256), so the run's final save is checkpoint-1590. No ambiguity about which weights are "the adapter".

Path correction. Section 13.2 records the M3-C adapter directory as m3c-full-6d20f38b53eb (the run-name hash). The directory actually written on Drive by the completed M3-C2 run is m3c-full-a6796d97783f0fec (the manifest-SHA prefix). Use the latter; 13.2's path does not exist on Drive.

14.2 Reproduction — obtaining the adapter

# On Colab, with Drive mounted (adapter is written straight to Drive by training):
ls /content/drive/MyDrive/SatQueryAI/adapters/m3c-full-a6796d97783f0fec

# To a local machine, into the git-ignored models/adapters/ tree:
mkdir -p models/adapters
# Option A — from a mounted Drive / rclone remote:
rclone copy "gdrive:SatQueryAI/adapters/m3c-full-a6796d97783f0fec" \
  "models/adapters/m3c-full-a6796d97783f0fec" --progress
# Option B — Google Drive web UI: download the folder, unzip into models/adapters/

# Integrity check (no torch required) -- hash + LoRA tensor count:
python scripts/verify_adapter_files.py models/adapters/m3c-full-a6796d97783f0fec/checkpoint-1590
# expect: sha256 3b625d63ed4f28640a6fb14072486305...
#         696 tensors, 37,152,768 parameters

Only adapter_config.json + adapter_model.safetensors are needed for inference; the optimizer/scheduler/RNG state in checkpoint-1590/ is required only to resume training and can be dropped for a deployment copy (saves 298 MB).

14.3 Evaluation harness (new)

File Purpose
eval/scoring.py Answer normalisation, per-family grading, error taxonomy. No torch — unit-testable on the laptop.
eval/run_adaptation_eval.py Base vs adapted on the TEST split: accuracy per question kind, per-CORINE-class metrics, error breakdown. Streams to JSONL and resumes.
eval/before_after.py Eight fixed prompts, base vs adapted, side-by-side into the evidence doc.
eval/docwrite.py Splices generated tables into marked blocks so hand-written prose survives regeneration.
tests/unit/test_adaptation_scoring.py 28 CPU tests, all passing.
docs/adaptation_evidence.md The write-up, for a reader who has not seen the code.

Design decisions worth recording:

  1. No headline aggregate accuracy. 7 of 13 question kinds are yes/no and the polarity leans 1.37:1 toward "no", so an aggregate flatters a model that learned only the majority answer. Everything is per question kind; a macro mean over kinds is the only summary reported.
  2. Base = disable_adapter() on the same loaded weights, not a second from_pretrained. Mathematically identical to the stock model, guarantees identical dtype/device/decoding on both sides, and halves VRAM.
  3. Over-listing is separated from hallucination in the error taxonomy — these are different problems and an accuracy number cannot distinguish them.

14.4 The overfitting finding, with a correction

Eval loss bottomed at epoch 2 and rose afterwards:

Step Epoch Eval loss
500 0.94 0.32608
1,000 1.89 0.30228
1,060 2.00 0.30211 ← best
1,500 2.83 0.30609
1,590 3.00 0.30574 (accepted)

The stop rule fired correctly and epoch 4 was rightly not launched.

However — the quoted final train loss of 0.06044 is an artefact, not a measurement. M3-C2 resumed from checkpoint-1060 after a 31.6-minute gap (elapsed_s resets at step 1070 in training_log.jsonl) and ran epoch 3 only. HuggingFace's end-of-run train_loss divides loss accumulated since resume by total steps:

sum(epoch-3 logged losses) × 10 / 1590 = 0.06044   ← reproduces the reported value exactly

Real per-epoch training loss (mean of logged step losses):

Epoch Mean Min Max
1 0.4084 0.1650 2.6218
2 0.1953 0.1402 0.2591
3 0.1813 0.1052 0.2479

The lowest single logged loss in the whole run is 0.1052. So the honest picture is saturation with a mild overfitting signal (train −7 % relative, eval +1.2 % over the final epoch), not a 0.30 → 0.06 train/eval collapse. The decision to stop is unchanged; the justification is diminishing returns with the eval curve turning, and that is what the write-up says.

14.5 Test suite after M3-D

python -m pytest -q → 109 passed, 6 skipped (81 + 28 new scoring tests). All 6 skips are the pre-existing architectural allowances in tests/test_backend_abstraction.py (modules permitted to import ML frameworks directly), not failures. No regression.

Caveat stated plainly: the suite is CPU-safe and mocks the model backend, so it verifies the harness, the tool contracts and the scoring rules — it does not exercise the adapted weights. Adapter behaviour is covered by the two eval scripts, which require a GPU session.

14.6 Status

  • Adapter downloaded, verified, and documented — done.
  • Eval harness written and unit-tested (28/28) — done.
  • Evidence doc written, including the overfitting correction — done.
  • Pending a GPU session: the base-vs-adapted numbers themselves. Both scripts splice their output into docs/adaptation_evidence.md on completion. Commands are in section 10 of that document.

15. M4 — Serving backend, captioning + grounding (TODO)

Implemented 2026-08-26 against the base model so the M3-C adapter can be swapped in later with no code change.

15.1 Serving abstraction (models/serving/)

File Purpose
models/serving/base.py InferenceBackend ABC
models/serving/local.py In-process backend, wraps models.vlm.registry
models/serving/remote.py HTTP client for the Colab FastAPI server
models/serving/factory.py get_backend() selected by MODEL_BACKEND env
models/serving/colab_api.py FastAPI server (/health, /infer, 1–2 images)
models/serving/readme.md Usage and env-var reference

Selection is purely configuration:

export MODEL_BACKEND=local        # laptop / judge machine
export MODEL_BACKEND=remote       # HTTP to Colab
export SATQUERY_REMOTE_URL=https://abc123.trycloudflare.com
export ADAPTER_PATH=/content/drive/MyDrive/SatQueryAI/adapters/m3c-full-6d20f38b53eb

colab_api.py prints the exact cloudflared command and the exact desktop env vars on startup.

15.2 New tools (agent/tools/)

  • agent/tools/captioning.py — SceneCaptioningTool, task captioning.
  • agent/tools/grounding.py — TextGuidedGroundingTool, task grounding.

Both conform to agent/tools/base.py Tool/ToolResult/ToolSpec. Both call the model only through models.serving.get_backend(). Neither module imports transformers/torch.

Grounding specifics:

  • Parses Qwen2.5-VL native <|box_start|>(x1,y1),(x2,y2)<|box_end|> boxes via regex, with JSON [x1,y1,x2,y2] / {"bbox": [...]} fallback.
  • Rescales 0–100 coordinates to original image pixels.
  • Renders overlay PNG into reports/grounding/.
  • Absent object returns ok=False, code=object_not_found, confidence ≤ 0.10.
  • Confidence is derived from parse success, box count, area sanity, refusal detection, and the rs_adapted flag; the heuristic is documented in the module docstring.

15.3 Local unit tests (CPU-safe, 25 tests)

python -m pytest tests/unit/test_captioning_tool.py tests/unit/test_grounding_tool.py -v

Result: 25 passed.

Full suite: 81 passed, 6 skipped.

15.4 Real-image VRSBench smoke test

Executed 2026-08-26 on the connected A100-80GB Colab runtime against the base model (ADAPTER_PATH unset). VRSBench images are nested under images/Images_val/, so the script now searches recursively.

Command used:

MODEL_BACKEND=local ADAPTER_PATH= \
  python scripts/test_m4_vrsbench.py \
    --data-root /content/data --n 5 \
    --report-dir /content/drive/MyDrive/SatQueryAI/reports/m4

Summary JSON: /content/drive/MyDrive/SatQueryAI/reports/m4/m4_vrsbench_summary_20260826_065711.json

Image Query GT box (norm) Caption conf Grounding conf Boxes Overlay path
07247_0000.png The vehicle located at the lower-left corner of the image. (0.07, 0.94, 0.16, 0.98) 0.25 0.10 (not found) 0 —
P0936_0019.png The harbor on the left side of the image. (0.10, 0.39, 0.24, 0.83) 0.25 0.10 (not found) 0 —
P0179_0049.png The plane is located near the bottom-middle of the image. (0.43, 0.67, 0.61, 0.86) 0.25 0.70 1 /content/drive/MyDrive/SatQueryAI/reports/m4/grounding/grounding_The_plane_is_located_near_the__20260826_065700.png
08579_0000.png The dark-colored vehicle parked at the bottom right of the image. (0.95, 0.70, 1.00, 0.82) 0.70 0.70 1 /content/drive/MyDrive/SatQueryAI/reports/m4/grounding/grounding_The_dark-colored_vehicle_parke_20260826_065706.png
P1471_0090.png The small vehicle positioned on the middle-left side of the image. (0.12, 0.61, 0.17, 0.70) 0.25 0.10 (not found) 0 —

Observations on the base model:

  • Captioning produces coherent scene descriptions on every image.
  • Grounding succeeds on 2/5 queries where the object is visually salient (plane, dark-colored vehicle). It returns explicit ok=False, confidence=0.10 on the other three, exactly the contract.
  • The two overlays are saved on Google Drive under reports/m4/grounding/.

15.5 Next step

M4 is complete. With the M3-C adapter saved to Drive, the next experiment is re-running the same smoke test with ADAPTER_PATH pointing at the adapted model and comparing caption/grounding confidence and box accuracy. That is an M3-D / M9 evaluation activity, not part of M4.


16. M6 — optical-SAR joint analysis (COMPLETE)

agent/tools/cross_modal.py — OpticalSARFusionTool, over the 1,000 co-registered Sentinel-2 / Sentinel-1 BigEarthNet v2 patches (Austria, 600 train / 200 validation / 200 test, selected deterministically by scripts/extract_ben_patches.py).

16.1 Quantitative branch: TRAINED, NOT DOWNGRADED

The milestone allowed a downgrade to NDWI/NDBI + SAR thresholding if the fusion head would not converge in 20 minutes on an A100. No downgrade was needed. The head converged on an L4 in 2.0 seconds, well inside budget:

metric built-up water
val average precision 0.766 0.811
val accuracy (calibrated) 0.690 0.865
val false-positive rate (calibrated) 0.191 0.055
fitted threshold 0.86 0.55

val AP mean 0.7885; 600 train / 200 validation patches; budget_exhausted: false. Artefacts: models/fusion/fusion_head.pt and training_report.json.

The head is a 1x1-convolution stack over the 12 stacked bands (S2 B02-B12 incl. NIR/SWIR + S1 VV/VH). Because a 1x1 conv is a per-pixel MLP, it trains on BigEarthNet's patch-level labels (per-patch band means) and applies unchanged at every pixel to produce a spatial map — the same weights in both directions.

The NDWI/NDBI + SAR threshold path is retained and still exercised: it runs whenever no trained head is supplied, and the tool then reports method="spectral_index_downgrade" in its output plus a QUANTITATIVE BRANCH DOWNGRADED warning, so the trace always says which branch produced the numbers.

16.2 CRITICAL TEST — SAR branch is live (3/3)

Each sample was answered three ways. both differed from optical only on 3 of 3 samples, so the SAR branch is not a rename of the optical one:

sample optical only both (fused)
UP_27_55 "no clear indication of built-up areas" "contains both built-up areas and water"
UP_27_61 "no clear indication of built-up areas" "contains both built-up areas and water"
UP_28_58 "no clear indication of built-up areas" "contains both built-up areas and water"

On every sample the optical-only answer denies built-up and the fused answer asserts it — the change can only have come from the radar input. The tool computes this itself (ablate=True → sar_contributes) rather than leaving it to an external script, and warns SAR BRANCH INERT with confidence capped at 0.50 if both ever collapses onto optical only.

16.3 Findings from running on real data

Three defects that the synthetic tests could not reach, all fixed and regression-tested:

  1. The downgrade path ignored SAR. Multiplying two hard-clipped terms saturates at 0 and 1, so wherever the optical index was already extreme the SAR factor changed nothing — exactly the "SAR is decorative" failure this milestone tests for. Now a weighted sum of soft logistic scores (OPTICAL_WEIGHT 0.6 / SAR_WEIGHT 0.4).
  2. A fixed 0.5 threshold was the wrong operating point. The BCE pos_weight that handles class imbalance inflates probabilities: 48% of built-up negatives and 28% of water negatives scored above 0.5. Thresholds are now fitted on validation data and stored with the weights. Fitting by F1 first made it worse (0.22 threshold, 95% FPR) because F1 ignores true negatives; calibration uses Youden's J, which prices false positives in.
  3. Presence was decided by pixel count, an uncalibrated rule. It marked every class present in every scene. The verdict now comes from the scene-level score the head was actually trained and calibrated on; coverage is still reported as spatial evidence, labelled via presence_basis.

16.4 Honest accuracy note

On the 3 test scenes the presence verdicts are 4/6 correct. The two errors (built-up false positive on UP_27_55, water false negative on UP_27_61) are consistent with the head's measured validation rates above, not a wiring fault. This is a small head trained on 600 patches of one country; it is reported as evidence with its error rates, not as a solved detector.

16.5 Evidence

  • Overlays (optical | SAR | built-up vs water mask): reports/m6/cross_modal/*.png
  • Run summary incl. all three ablation answers: reports/m6/m6_summary_*.json
  • Tests: tests/unit/test_cross_modal_tool.py (24 tests)

17. M7 — agentic controller, routing, validation, execution trace (COMPLETE)

agent/planner/ and agent/validators/. One entry point, AgentController.run, one output, an ExecutionTrace. Every stage writes into the trace as it goes, so a rejected run still produces a complete record instead of an exception.

17.1 Routing

classifier.py is hybrid: keyword rules first (fast, deterministic, and they carry the five PS queries), model fallback only for genuinely ambiguous wording. Image count and detected modality are evidence for routing, not just for validation — a query that clearly asks for change analysis but supplies one image is classified unsupported with intended_task=change_analysis, so the trace carries the specific typed error rather than a vague "unsupported".

17.2 Validation

validators/inputs.py, checked cheapest-and-most-decisive first, and every check is recorded whether or not it fails: existence and decodability, image count vs. task requirement, format and georeferencing, modality detection with the evidence that supports it, and pair geometry (CRS, footprint, shape). Typed errors: TooFewImages, ModalityMismatch, UnregisteredFormat, NotCoRegistered, CorruptImage. validate() never raises.

17.3 Confidence: weakest-link, and why

overall = min(task_confidence, *step_confidences) minus penalties (failed step −0.20, benchmark-mode input −0.05), capped at 0.70 when the backend is not RS-adapted. An average would let a 0.95 classification hide a 0.35 answer; the minimum cannot. The UI names the weakest component and its value.

17.4 The one real chain

A change query that also asks where runs ChangeAnalysis and then Grounding on the change map step 0 produced — a genuine data dependency, recorded in the plan as consumes: step_0.change_map_evidence.

17.5 Results — all 7 scenarios, against the adapted model

Run 2026-08-27 through scripts/test_m7_controller.py --backend remote against Qwen2.5-VL-3B + checkpoint-1590 on an NVIDIA L4 (rs_adapted=true).

# Scenario Routed to Status Confidence
1 captioning, 1 optical captioning success 0.85
2 grounding, 1 optical grounding success 0.80
3 change + where, bi-temporal change_analysis -> chain success (2 steps) 0.35
4 optical + SAR cross_modal success 0.80
5 direction, bi-temporal change_analysis success 0.35
6 change query, ONE image unsupported failed 0.00 (too_few_images)
7 mismatched CRS pair change_analysis failed 0.00 (not_co_registered)

7/7 routed correctly. Queries 3 and 5 score 0.35 because the disagreement cross-check fired: the model says "no significant change" while the change map marks 3.5% of the scene. That is the design working, and the language answer being wrong.


18. M8 — interactive web application (COMPLETE)

app/frontend/app.py. Streamlit, because the deliverable is an evidence panel, not a bespoke SPA.

  • Upload zone showing modality, bands, size, dtype and CRS before submit, with a pair-compatibility verdict and a live routing preview.
  • The five PS sample queries as one-click buttons, verbatim.
  • Answer, inputs beside generated visual evidence, confidence as a number with its basis stated.
  • Execution trace panel as the primary surface: trace id, timestamp, classified task with method, model and adapter ids plus adapter sha256, the plan including what each step consumes, then one row per step with tool, status, duration, per-step confidence, model, adapter, bound params and output.
  • trace.json and report.pdf downloads, working on rejected runs too.
  • Sidebar connection indicator: a dead cloudflared tunnel is shown as unreachable, never as a hang.

Screenshots of all nine states in docs/images/ui_*.png, captured against the adapted model on the L4.


19. M9 — hardening, docs, and four real defects (COMPLETE)

19.1 Four defects found by running the PS queries against the real model

The stub backend had been hiding all four. Each was found only by pointing the controller at Qwen2.5-VL-3B + adapter on a GPU.

  1. patch_embeddings read a quarter of the scene as if it were the whole thing. transformers 4.x returned the vision tower's merged output (one 2048-d token per 2x2 cell); 5.x returns pre-merger tokens (four 1280-d tokens per cell). Taking the first h*w of those sampled a corner in block order, and raw cosine similarity between those features is 0.96 even for a black image versus a white one — so the change map was noise. Fixed by running the merger explicitly. Verified by blanking a known ninth of a scene and confirming those cells, and only those, go hot.
  2. Grounding never parsed a box. The adapter emits a bare pixel quadruple (153,21,224,91), not the base model's <|box_start|> tokens, so every grounding query returned object_not_found. Confirmed against ground truth: a fixture settlement at (0.60, 0.08)-(0.88, 0.36) comes back as (0.598, 0.082, 0.875, 0.355) once divided by the image dimensions.
  3. Grounding was handed the whole query as the object name, so the PS's own "Highlight the water body referred to in the query." asked the model to find an object of that name. Now the noun phrase is extracted, and a query that yields nothing is retried against the adapter's CORINE vocabulary — it localises river where it answers not-found to water body. Every phrasing tried is recorded in the trace.
  4. Single-image VQA — PS capability 1 — was broken on the remote backend. It predates models.serving and still called generate(image, prompt, GenerationConfig) instead of generate(images: list, prompt, params: dict). The local backend tolerated it; the remote one failed with 'Image' object is not iterable. The test fake had been written against the old interface, so it asserted the broken call was correct, and the M7 harness never covered plain VQA. Both gaps closed.

19.2 Change map: two channels, fused by agreement

Patch-cosine distance is semantic and illumination-robust but only as fine as the patch grid; pixel differencing is sharp but fooled by radiometric differences between dates. Their failure modes are opposite, so the map is their geometric mean — a cell survives only where both agree. Measured on the co-registered fixture pair (change = 6.0% of pixels):

Method IoU F1
patch cosine alone 0.13 0.23
pixel difference alone 0.66 0.79
fused (shipped) 0.55 0.71

Pixel differencing wins there because that fixture has no radiometric difference between dates at all — precisely the case it is best at and the case real imagery does not give you. The conservative agreement rule ships, and every channel's statistics go into the trace so the choice stays auditable. On real imagery with illumination differences the ranking may invert; that is unmeasured.

19.3 Imaging hardened for real GeoTIFF

agent/imaging.py: RGB bands chosen from colour interpretation, then band descriptions, then a documented count heuristic (a 12-band Sentinel-2 stack no longer renders as coastal/blue/green); nodata masks excluded from stretch statistics and rendered black; rasters above 2048 px read decimated; SAR converted to dB and normalised on median/IQR rather than min/max, because speckle is multiplicative and heavy-tailed; reproject_to_match/align_pair for pairs.

19.4 Tests

253 passing, 11 skipped (from 179/11). tests/unit/test_m9_hardening.py adds 68 written on the premise that a judge uploads a real Cartosat/RISAT GeoTIFF: arbitrary band counts and dtypes, nodata, decimated reads, SAR dB conversion, every validator error path, per-tool input validation, registry/validator agreement, controller routing for all five PS queries, trace JSON round-trip and PS-named fields, and typed failure (never a crash) on empty, corrupt and truncated files. tests/fixtures/synthetic.py generates every fixture deterministically, so no imagery is committed.

19.5 Repository hygiene

.git is 28 MB; the largest blob ever committed is a 1.85 MB screenshot. No weights, datasets, .env or secret-shaped strings anywhere in history. The duplicate ## 13. M3-C heading is resolved (§13 and §13.A).

19.6 Clean-clone verification

Cloned fresh from GitHub into a temp directory and followed only the README: pip install -r requirements.txt succeeded, pytest gave 253 passed / 11 skipped, the fixture generator produced all 13 files, the Streamlit app booted and served HTTP 200, the controller harness passed 7/7 on the stub and 7/7 against the live remote adapted model, and make_architecture_diagram.py regenerated the committed PNG byte-identically. Nothing in the README needed correcting. Following it leaves no untracked files.

19.7 Hardware incident

Mid-milestone the F: USB SSD disconnected entirely — the documented Risk 1 in section 10, recurring. Nothing was lost (the volume remounted and git fsck was clean), but one commit had been sitting unpushed for about an hour. Practice changed for the rest of the session: push after every commit, not at the end.


20. M11 — night-before-judging connection audit (COMPLETE)

A user report ("pasted a Colab tunnel URL, got an error, could not proceed") led to a full audit of the connection path, the front end, and a real end-to-end run of every PS-mandatory query. Full narrative and evidence: docs/PS_COMPLIANCE.md §9 and its "2026-09-02, later the same session" addendum. Summary here for the milestone log:

Real bugs found and fixed (each its own commit, see git log):

  1. /api/health reported the same state for a working model and the stub backend — the actual reported failure, not CORS. d2a39fd.
  2. Neither shell launcher searched adapter/, the path actually committed via git-lfs. d2a39fd, db81a7c.
  3. CORS tightened to an allowlist + optional token; the allowlist itself then had to grow twice (a second production URL, below). d2a39fd, e6e840c.
  4. A trailing slash on the pasted API base 404'd every request. 0537d46.
  5. The LoRA adapter silently failed to load on a fresh Colab install — peft needs torchao>=0.16.0, never pinned, so Colab's older preinstalled version stood and the query still answered successfully off base weights. A real PS §1 risk, invisible unless you read the trace warnings. 962edd4. Pin committed; not yet re-verified against a second live Colab run (kernel access ran out this session).
  6. agent/tools/grounding.py crashed on any query its spectral checker recognised — warnings.append() before warnings was defined. PS sample query #2 ("highlight the water body") hit this on every run. Found by actually running all 5 PS queries live, not by a test. e4fe52d.
  7. demo_assets/*.tif is gitignored, so a fresh clone (Colab included) has no sample imagery until scripts/fetch_real_demo_assets.py runs — never wired into a launcher. Worked around by uploading a local image instead; not fixed at the code level yet.

Deployed: a second production URL, https://satquery-ai-ps26167.netlify.app, went live via an already-authenticated Netlify connection after wrangler login's interactive OAuth wasn't completed in time. Carries the mission-control redesign (already committed, never live before tonight — the original dfbc41d redesign predates this milestone), the location-fix, and every fix above.

Verified live on that URL: location search (real STAC timeline, a real place), all 5 PS sample queries (one fix required, #6 above), one rejection case (not_co_registered, typed, no crash). Test suite: 272 passed, 11 skipped, 1 pre-existing unrelated failure (test_largest_region_finds_the_water_square, not investigated).

Not completed this session, and why: VRSBench/RSVQA/CDVQA benchmark numbers (Colab kernel access ran out before a clean download window opened — see CLAUDE.md's environment notes for the single-kernel constraint); a second live Colab run to confirm fix #5; docs/videos/connect_colab.mp4 re-recorded against tonight's fixes (still the pre-session recording). All three need Colab kernel time this session didn't get. CLAUDE.md added specifically so the next session (or a different account) doesn't have to rediscover any of this.