An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries
- Try it now
- Problem Statement β PS 26167
- Quickstart
- The five problem-statement queries
- Architecture
- Screenshots & Demo
- Tech Stack
- Model & Adaptation
- Evaluation & Results
- Use Cases
- Impact
- Future Scope
- PS requirement β file mapping
- Known limitations
- Read more
- Team
No installation required.
Point it at a backend from the connection bar at the top (see Quickstart). Backup mirror (may lag a build or two behind): https://satquery-ai.pages.dev.
The live page must talk to an
https://backend. A browser blocks plain-HTTP requests from an HTTPS page, sohttp://localhost:8000will not work from the hosted app (it works fine when you run the front end locally).run_backend.pyprints exactly the right kind of URL β ahttps://β¦.trycloudflare.comtunnel.
| Problem Statement ID | 26167 |
| Title | SatQuery AI β An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries |
| Organization | Indian Space Research Organisation (ISRO) |
| Department | Department of Space / ISRO |
| Category | Software |
| Theme | Space Technology |
In short: most remote-sensing AI tools are single-task, single-modality, and assume a GIS-literate user. PS 26167 asks for an agentic assistant that takes a plain-English question plus one or two satellite images, figures out which specialist task is being asked for, checks whether the images it was given can actually answer that question, runs the right tool(s) β including joint reasoning over optical+SAR pairs and multi-temporal pairs, not just single images β and returns an answer backed by visual evidence and an auditable execution trace. A generic, un-adapted VLM is explicitly disqualified.
| # | PS mandatory requirement | Our solution |
|---|---|---|
| 1 | Remote-sensing adaptation β at least one vision/VLM component fine-tuned on BigEarthNet.txt or equivalent open data | Qwen2.5-VL-3B + a custom LoRA adapter trained on BigEarthNet β see Model & Adaptation |
| 2 | Single-image baseline β VQA mandatory, plus captioning or grounding | vqa.py (free-form Q&A) and both captioning.py + grounding.py (we built both, not just one) |
| 3 | Multi-image change analysis β change description/VQA from a bi-temporal pair, spatial change map where possible | change.py β change-VQA and a rendered before/after/change-overlay map |
| 4 | Cross-modal pair analysis β complementary information from co-registered optical + SAR | cross_modal.py β VLM reasoning + a trained fusion head, tested 3 ways (optical-only / SAR-only / both) to prove SAR actually changes the answer |
| 5 | Agentic orchestration β classify, validate inputs, select/sequence tools, combine outputs, auditable trace | planner/classifier.py β planner/controller.py β planner/trace.py β never raises, always returns a complete trace |
π Full official problem statement text (click to expand)
Background. Remote-sensing imagery is widely used for agricultural monitoring, disaster management, urban planning, forest monitoring, water-resource assessment, infrastructure mapping, and environmental analysis. However, most existing remote-sensing AI solutions are developed as isolated applications for a single predefined task, such as land-cover classification, object detection, visual question answering, or change detection. These systems often require users to understand satellite-data characteristics, GIS workflows, model selection, and task-specific parameters. Consequently, non-expert users may find it difficult to obtain meaningful information from satellite imagery through simple natural-language queries.
Many operational remote-sensing questions cannot always be answered reliably using a single optical image. Relevant information may be distributed across paired or multiple observations acquired at different times or by different sensors. Optical and multispectral imagery provides spectral and contextual information, whereas synthetic aperture radar (SAR) provides complementary structural information and supports day-and-night acquisition through cloud cover. Multitemporal image pairs are required to identify and interpret changes over time, while co-registered opticalβSAR pairs can provide more complete and reliable information than either modality alone.
A general-purpose large language model (LLM) or vision-language model (VLM) cannot be expected to perform these specialised tasks reliably without adaptation to remote-sensing imagery, sensor characteristics, and domain-specific terminology. The proposed solution must therefore include remote-sensing fine-tuning or domain adaptation and may employ multiple specialised models for different tasks. BigEarthNet.txt will serve as the primary dataset for adapting imageβtext representations to multisensor remote-sensing data. VRSBench and RSVQA will be used to evaluate single-image captioning, grounding, and visual question answering, while CDVQA will be used to evaluate multitemporal change-based visual question answering.
The novelty of SatQuery AI lies in its agentic, query-driven framework. Instead of applying a single generic VLM, the system selects and executes suitable remote-sensing specialist models, validates inputs, combines their outputs, and returns an evidence-grounded response.
Description. The objective is to develop SatQuery AI, a software-based agentic vision-language assistant for analysing single and paired remote-sensing images through natural-language queries. Single-image understanding is a mandatory baseline, while the principal focus is joint reasoning over paired cross-modal and multitemporal imagery.
Defined Input Scope
- Single image: one optical/multispectral or SAR image for captioning, visual question answering, and text-guided region grounding.
- Cross-modal pair: co-registered optical/multispectral and SAR images of the same geographic area for joint information extraction and cross-modal analysis.
- Bi-temporal pair: two spatially corresponding images of the same geographic area acquired at different times for change detection, change description, and change-based visual question answering.
- Supported formats: GeoTIFF or TIFF for geospatial imagery. PNG and JPEG inputs may be accepted only for the prescribed public benchmark datasets.
Mandatory Functional Scope
- Remote-sensing adaptation: at least one visual or vision-language component must be fine-tuned or otherwise adapted using BigEarthNet.txt or any open-source training data.
- Single-image baseline: visual question answering shall be mandatory. Each solution must additionally implement either captioning/scene description or text-guided region grounding.
- Multi-image change analysis: change description or change-based visual question answering from a bi-temporal image pair shall be mandatory. A spatial change map may also be generated where reference masks are available.
- Cross-modal pair analysis: the system must extract complementary information from a co-registered optical/multispectral and SAR image pair.
- Agentic orchestration: the system must automatically select, sequence, and execute the appropriate specialist models or tools according to the query and input configuration.
Representative Queries
- "Describe the land-cover and major objects visible in this image."
- "Highlight the water body referred to in the query."
- "What changed between these two dates, and where did the change occur?"
- "Use the optical and SAR images together to identify built-up and water-covered regions."
- "Has the built-up area increased, decreased, or remained unchanged?"
Agentic Model and Tool Orchestration. The system may use multiple specialised components, such as a remote-sensing VQA or captioning model, a grounding model, a change-understanding or change-VQA model, and an opticalβSAR fusion or information-extraction model. The controller must:
- interpret the query and classify the requested task;
- check the number, modality, format, metadata, and compatibility of the input images;
- select one or more models or tools from a predefined registry;
- configure only permitted task parameters and execute the selected workflow;
- combine textual and spatial outputs, estimate confidence, and return visual evidence; and
- provide an auditable execution summary containing the selected task, model/tool names, and key parameters.
The controller may perform internal task planning; however, only the observable execution trace, including the selected task, models or tools, permitted parameters, and outputs will be evaluated. Internal reasoning text is neither required nor evaluated.
Expected Solution. An interactive GUI or web application with an agentic remote-sensing AI backend. It should accept supported image inputs and natural-language queries, select the appropriate specialist workflow, and return evidence-grounded textual and visual results. The solution should include:
- input upload and compatibility checking;
- a remote-sensing-adapted vision-language component;
- specialist tools for VQA, captioning or grounding, change understanding, and opticalβSAR analysis;
- an agentic controller for task routing, tool execution, and output integration;
- visual evidence, confidence information, execution summaries, and downloadable reports.
Each solution must demonstrate single-image VQA, one additional single-image task, multitemporal change understanding, opticalβSAR paired-image analysis, and agentic model/tool orchestration. A generic LLM or VLM without remote-sensing adaptation will not satisfy the requirements.
Deliverables. An interactive GUI or web application with an agentic remote-sensing AI backend, codes and models, including test and demonstration.
Implementation Scope. The system shall support single optical/multispectral or SAR images, co-registered opticalβSAR pairs, and bi-temporal pairs in GeoTIFF/TIFF or approved benchmark formats. It must perform single-image VQA, one additional single-image task, change analysis, opticalβSAR joint analysis, and agentic model/tool selection through an interactive GUI or web application.
Evaluation / Judging Criteria. Final evaluation will use prescribed public benchmark test subsets and an ISRO/SAC evaluation dataset. Scores will be normalised before combining different metrics. Public benchmarks will be evaluated using the prescribed test splits. The ISRO/SAC evaluation set will contain pre-georeferenced and co-registered Cartosat-2S optical and RISAT SAR image pairs, with task-specific reference answers, labels, bounding boxes, or masks, as applicable. Evaluation annotations will not be disclosed to participating teams.
Dataset Link. BigEarthNet.txt β primary dataset for remote-sensing adaptation using co-registered Sentinel-1 SAR, Sentinel-2 multispectral imagery, and diverse text annotations. All datasets are available online, open source. Public evaluation benchmarks: VRSBench, RSVQA (single-image captioning/grounding/VQA), and CDVQA (multitemporal change-based VQA).
Pick the tier that matches what you have. All three end with the same step: paste the printed URL into the connection bar above and click "Use & check".
|
Colab β free, no GPU needed !git clone https://github.com/KING-OF-FLAME/satquery-ai.git
%cd satquery-ai
!python run_backend.pyRuntime β Change runtime type β GPU, before running the cell. |
Your own GPU β 8 GB+ VRAM git clone https://github.com/KING-OF-FLAME/satquery-ai.git
cd satquery-ai
python run_backend.py
|
No GPU β UI only, fake answers python run_backend.py --backend stubAnswers are canned and clearly
labelled |
Full walkthrough, troubleshooting, and every flag: docs/SETUP.md Β·
docs/TROUBLESHOOTING.md.
| Query (verbatim from the PS) | Capability |
|---|---|
| "Describe the land-cover and major objects visible in this image." | Single-image captioning |
| "Highlight the water body referred to in the query." | Text-guided grounding |
| "What changed between these two dates, and where did the change occur?" | Bi-temporal change analysis |
| "Use the optical and SAR images together to identify built-up and water-covered regions." | Cross-modal fusion |
| "Has the built-up area increased, decreased, or remained unchanged?" | Change direction |
Click any of them as one-click sample buttons in the app, or type your own.
Generated by scripts/make_architecture_diagram.py
and committed as a real file, so it can be regenerated instead of redrawn.
ASCII fallback (for viewers that do not render images)
Streamlit web app Scripted entry points Inputs
app/frontend/app.py scripts/test_m7_controller.py GeoTIFF/TIFF (georeferenced)
upload Β· query Β· trace eval/ Β· notebooks/ PNG/JPEG (benchmark mode)
downloads trace.json+pdf same controller, no UI BigEarthNet patch directory
| | |
+------------------------------+------------------------------+
v
+--------------------------------------------------------------------------+
| AGENTIC CONTROLLER agent/planner/controller.py |
| 1 classify -> 2 validate -> 3 plan -> 4 execute -> 5 combine |
| classifier.py validators/ bind only per-step weakest-link |
| inputs.py permitted timing, confidence |
| params never raises |
+--------------------------------------------------------------------------+
| | | | |
v v v v v
single_image scene_ text_guided_ change_ optical_sar_
_vqa captioning grounding analysis fusion
vqa.py captioning.py grounding.py change.py cross_modal.py
1 optical 1 optical 1 image -> bbox 2 dates -> optical + SAR
^ change map
| |
+---------------------+
chain: ground the change map
| | | | |
+------------+------------------+-----------------+--------------+
v
INFERENCE BACKEND models/serving/ (selected by MODEL_BACKEND)
LocalBackend RemoteBackend StubBackend
local.py remote.py stub.py
in-process HTTP -> SATQUERY_REMOTE_URL no weights
CPU: minutes/query /health /infer /embeddings UI tests only
v
Qwen2.5-VL-3B-Instruct + LoRA ExecutionTrace agent/planner/trace.py
checkpoint-1590 classified_task Β· task_confidence Β· plan
BigEarthNet v2, 11,763 records steps[tool, model_id, adapter_id, params,
duration_ms, status, output_summary]
Graceful failure: a run rejected at validation (too_few_images Β·
modality_mismatch Β· not_co_registered Β· unregistered_format Β· corrupt_image)
still returns a complete trace: zero steps executed, confidence 0.00.
Two design decisions worth calling out. The trace is written as the run
happens, not assembled at the end β a run that is rejected, or whose tool crashes,
still produces a complete record; AgentController.run never raises. Confidence is
weakest-link, not an average β overall = min(task_confidence, *step_confidences),
so a confident classification can never hide an unreliable answer.
Full rationale, repo layout, and the complete PS-requirement β file table:
docs/ARCHITECTURE.md.
All screenshots are real captures against the RS-adapted model on a Colab GPU β
not mockups. Full gallery (all 5 queries, both failure cases, videos):
docs/GALLERY.md.
|
π¨ Frontend Static-exported Next.js app, no server components β every backend call is made from the browser. |
βοΈ Backend Job-based HTTP API wrapping the same agent, validators and trace format the CLI uses. |
π§ AI / ML Qwen2.5-VL-3B + a custom LoRA adapter β see below. |
βοΈ Data & Deploy No traditional database β run
state is in-memory per server
process; imagery, traces and
reports persist to the filesystem
( |
Full stack rationale (why each choice, not just what): docs/ARCHITECTURE.md.
|
Base model
Why adaptation was necessary The base model produces fluent prose but the wrong shape of answer for a
pipeline: asked to list land-cover classes it says "not possible to list
specific land-cover classes"; asked to localise a river it says |
LoRA adapter
|
Training data. BigEarthNet.txt / BigEarthNet v2 β the PS-named
dataset β co-registered Sentinel-2 (multispectral) + Sentinel-1 (SAR) with official
CORINE land-cover labels and caption/QA text annotations. 11,763 instruction
records built by build_manifest.py,
hash-verified (a6796d97β¦) so the exact training set of this adapter is auditable.
Why it stopped at epoch 3. Held-out loss rose between epoch 2 (0.3021) and
epoch 3 (0.3057) while train loss kept falling (0.30 β 0.06) β textbook overfitting on
a 600-patch split. A fourth epoch would score better on paper and answer worse in
practice, so checkpoint-1590 ships instead. Full account, including a correction to
an earlier headline number, in docs/adaptation_evidence.md.
How it was built, reproducibly, in four commands:
python training/geochat_adaptation/build_manifest.py --raw data/raw ... # deterministic manifest
python training/geochat_adaptation/train_lora.py --config .../m3b_smoke.yaml # 100-sample smoke test
python training/geochat_adaptation/train_lora.py --config .../m3c_full.yaml # the real run, ~30 min on an A100/L4
python training/geochat_adaptation/verify_adapter.py --adapter_path <out>/checkpoint-1590Full reproduction guide, exact hyperparameters, and the loss table:
docs/SETUP.md.
Every number below was produced by actually running the system β where a number does not exist yet, the table says so rather than leaving a blank that reads as a pass.
| Metric | Result | Source |
|---|---|---|
| Adapter effect (base vs. adapted, fixed prompts) | 7 / 7 prompts changed β base model factually wrong on SAR backscatter, adapted model correct | eval/fixture_before_after.py |
| Controller routing β all 5 PS queries + 2 failure gates | 7 / 7 correct, including the 2-tool changeβgrounding chain | scripts/test_m7_controller.py |
| Change-map IoU (fused channel) | 0.55 vs. pixel-exact ground truth | co-registered fixture pair |
| Cross-modal presence verdicts | 4 / 6 correct | fusion head, 600 BigEarthNet patches |
| Grounding spectral verification | 51% water-in-box vs. 4% in-scene (14Γ enrichment) | NDWI measured on Navi Mumbai scene |
| Test suite | 278 passed, 6 skipped (GPU/full-dataset only) | pytest tests/ -q |
Confidence is weakest-link, never an average β one unreliable step caps the whole
score (overall = min(step confidences), capped at 0.70 with no adapter, β0.05 for
benchmark-mode input). A number without its basis is decoration; the trace names which
component was weakest.
Honestly, what's not measured yet, and why. The PS names four evaluation
sources: VRSBench and RSVQA (single-image captioning/grounding/VQA),
CDVQA (multitemporal change VQA), and an undisclosed ISRO/SAC set
(Cartosat-2S optical + RISAT SAR). Formal accuracy on these prescribed splits has
not been run β the scoring harnesses exist and are unit-tested, but running them
needs the benchmark imagery itself (BigEarthNet's full split alone is ~118 GB at a
measured 0.7 MB/s β 25 hours on this connection). The ISRO/SAC evaluation is out
of reach for every team by design β its annotations are undisclosed to participants
before judging, exactly as the PS states. Full breakdown, including known limitations
stated plainly: docs/EVALUATION.md Β· docs/LIMITATIONS.md.
Straight from the PS's own list of who needs this β one row per domain, tied to the query type that actually answers it.
| Domain | Question a user asks | SatQuery AI capability |
|---|---|---|
| πΎ Agricultural monitoring | "What's the land-cover composition of this field?" | Single-image captioning / VQA |
| π Disaster management (floods) | "Has the water body grown since last month?" | Bi-temporal change analysis + change map |
| ποΈ Urban planning | "Where is the built-up area, and is it expanding?" | Change direction + opticalβSAR fusion |
| π² Forest monitoring | "Highlight the forested region in this scene." | Text-guided grounding |
| π§ Water-resource assessment | "Highlight the water body referred to in the query." | Grounding + spectral (NDWI) verification |
| ποΈ Infrastructure mapping | "What changed between these two dates, and where?" | Two-tool chain: change analysis β grounding |
| π°οΈ All-weather / night monitoring | "Use optical and SAR together to find built-up and water regions." | Cross-modal fusion β SAR sees through cloud cover and works in the dark |
| π Environmental analysis | Any of the above, on any place | Live location search β fetches real Sentinel-2 for any spot on Earth |
Every one of these runs through the same agentic controller β a user never has to know which model, tool, or GIS parameter answers their question.
- Removes the GIS-literacy barrier. The PS's own premise: non-expert users struggle to get answers from satellite imagery today because it demands knowing sensor characteristics, model selection, and task-specific parameters. SatQuery AI reduces that to a plain-English question and an image.
- Verifiable, not just plausible. Every answer ships with the exact task, tool, model + adapter id, parameters, and a confidence number whose weakest-link basis is named β decision-makers in disaster response, defence, or policy can check an answer rather than trust it blindly. This matters more than raw accuracy for high-stakes domains.
- All-weather, day-and-night coverage. By fusing optical with SAR, the system keeps working through cloud cover and at night β exactly when disaster monitoring (floods, cyclones) needs it most and optical-only tools go blind.
- One system, many missions. Agriculture, disaster response, urban planning, forestry, water resources, and infrastructure mapping are usually separate, single-purpose tools built by separate teams. One agentic controller here routes all of them, which is cheaper to build, maintain, and extend than five silos.
- A path to sovereign Indian EO tooling. Built around ISRO's own evaluation target (Cartosat-2S + RISAT); Sentinel-2/-1 stand in today only because ISRO/SAC annotations are withheld until judging, by design.
- Close the accuracy loop. Run the prescribed benchmarks β VRSBench, RSVQA, CDVQA, and the full BigEarthNet held-out split β once bandwidth allows, and publish real numbers instead of the fixture-based evidence used today.
- Evaluate on real ISRO/SAC imagery. Cartosat-2S optical + RISAT SAR are the actual target sensors; Sentinel-2/-1 are proxies. Swapping in real Indian satellite data is the single highest-value next step for PS compliance.
- Hard-negative supervision. Teach the adapter which CORINE classes are absent, not just which are present, to close the residual over-listing gap.
- Bi-temporal supervision at scale. The adapter saw CDVQA pairs only lightly; more change-focused training data is where the biggest accuracy gain likely sits.
- More sensors, more bands. Extend beyond Sentinel-2/-1 to hyperspectral, thermal, and additional SAR polarizations as they become available.
- Active learning from real usage. Let judges' and users' corrections feed back into future fine-tuning rounds instead of a single static adapter.
- Mobile and offline-first clients. A field-usable app for disaster responders and farmers without reliable connectivity, syncing traces when back online.
- Public API + third-party integration. Expose the controller as a stable API so GIS platforms, government dashboards, and other tools can call SatQuery AI as a reasoning layer rather than reimplementing it.
- Alerting, not just querying. Turn the existing change-alert thresholds into a standing watch β "tell me if this area changes" β instead of only answering on-demand.
| # | Mandatory requirement | Where it lives |
|---|---|---|
| 1 | Remote-sensing adaptation of a VLM | training/geochat_adaptation/train_lora.py, adapter checkpoint-1590, evidence in docs/adaptation_evidence.md |
| 2 | Single-image VQA | agent/tools/vqa.py |
| 2 | A second single-image task β captioning and grounding | agent/tools/captioning.py, agent/tools/grounding.py |
| 3 | Bi-temporal change analysis with a rendered change map | agent/tools/change.py |
| 4 | OpticalβSAR cross-modal analysis using both modalities | agent/tools/cross_modal.py, agent/fusion.py |
| 5 | Agentic routing of a free-text query to the right task | agent/planner/classifier.py, agent/planner/controller.py |
| 5 | Observable execution trace: task, tools, model/adapter ids, permitted params, outputs | agent/planner/trace.py |
| 6 | Input compatibility validation and graceful failure | agent/validators/inputs.py, typed errors in agent/errors.py |
| 6 | GeoTIFF support: arbitrary bands, dtypes, nodata, CRS | agent/imaging.py |
| 7 | Interactive web application | app/frontend/app.py, app/frontend-web/, app/webapi/ |
| 7 | Visual evidence + confidence displayed | Results in app/frontend-web/src/components/Results.tsx |
| 7 | Downloadable report (PDF) and trace (JSON) | app/frontend/report.py, build_pdf in app/webapi/server.py |
Full mapping including non-mandatory extras: docs/ARCHITECTURE.md.
Stated plainly, because a system that hides these is harder to trust than one that
does not. Full detail: docs/LIMITATIONS.md.
- Bi-temporal language answers are the weakest output β the system detects when its own words disagree with the measured change map and says so, capping confidence.
- No held-out benchmark accuracy yet on VRSBench / RSVQA / CDVQA / BigEarthNet (imagery bandwidth-bound, not code-bound β see Evaluation above).
- No Cartosat-2S / RISAT evaluation has been possible for any team β annotations are undisclosed before judging. Sentinel-2/-1 are used as the closest open proxies.
- Grounding depends on a fixed synonym list per concept; an object phrased outside that list may fail to localise even when visible.
docs/ARCHITECTURE.mdβ system diagram, design decisions, stack rationale, repo layout, PS requirement β file mappingdocs/GALLERY.mdβ screenshots, videos, and results for all five queries plus failure casesdocs/EVALUATION.mdβ base-vs-adapted comparison, measured metrics, what's not yet measured and why, confidence rule, test suitedocs/SETUP.mdβ full local setup, the complete Colab walkthrough, the full end-to-end usage guide, reproducing the LoRA fine-tunedocs/TROUBLESHOOTING.mdβ "unreachable" fixes, requirements tables, redeploying the front enddocs/LIMITATIONS.mdβ known limitations, stated plainly, and next stepsdocs/adaptation_evidence.mdβ the LoRA adaptation's own evidence documentdocs/DEMO.mdβ four-minute demo scriptdocs/PS_COMPLIANCE.mdβ problem-statement compliance checklistPROJECT_STATUS.mdβ full engineering log, every milestone
| Name |
|---|
| Yash Raj |
| Dhruv Devaliya |
| Tarak Dhone |
| Lakshay Vig |
| Devendra Rajpuhoit |
| Krishna Dhaker |
Smart India Hackathon 2026 Β· Problem Statement 26167 Β· Indian Space Research Organisation (ISRO)
π°οΈ Made for SIH 2026 Β· PS 26167 Β· Try SatQuery AI live β









