An interactive web prototype for SatQuery AI: an agentic vision-language assistant for multimodal remote-sensing image analysis.
- Single-image natural-language analysis
- Remote-sensing caption/VQA-style workflow
- Text-guided region grounding prototype
- Bi-temporal before/after change analysis with a real pixel-difference evidence map
- Optical + SAR joint-analysis prototype with visual fusion overlay
- Input validation
- Agentic task routing
- Specialist tool registry
- Observable execution trace
- Confidence score
- Downloadable JSON analysis report
- GeoTIFF/TIFF/PNG/JPEG input acceptance at the API layer
The build uses transparent image-processing baseline adapters for the demo. It does not falsely claim that these heuristics are a fine-tuned remote-sensing VLM. For the SIH evaluated build, replace the adapter functions in backend/analyzer.py with actual remote-sensing checkpoints/fine-tuned components and document the training/adaptation procedure.
The orchestration, API, UI, evidence system, and model-adapter interfaces can remain unchanged.
- Install Python 3.10+.
- Open Command Prompt/PowerShell in this folder.
- Create a virtual environment:
python -m venv .venv
.venv\Scripts\activate- Install dependencies:
pip install -r backend/requirements.txt- Start the server:
python backend/app.py- Open:
http://127.0.0.1:5000
For the fastest demo, use PNG/JPEG satellite-style images. The interface accepts TIFF/GeoTIFF extensions as well. If you need full multispectral GeoTIFF band handling, add Rasterio/GDAL preprocessing in load_visual().
Upload one image and ask:
Describe the land-cover and major objects visible in this image.
Show:
- task classification
- RS captioning/VQA specialist
- answer
- confidence
- evidence
Select Change Analysis, upload before/after images, ask:
What changed between these two dates, and where did the change occur?
Show:
- before
- after
- candidate change map
- execution trace
Select Optical + SAR, upload optical + SAR images, ask:
Use the optical and SAR images together to identify built-up and water-covered regions.
Show:
- optical evidence
- SAR evidence
- joint overlay
- fusion integrator
The clean replacement points are:
caption()→ remote-sensing caption/VLM checkpoint- VQA branch → remote-sensing VQA checkpoint
grounding()→ grounding/segmentation modelchange_analysis()→ trained bi-temporal change detector + change captionerfusion_analysis()→ optical-SAR fusion model
Keep the same JSON response contract:
{
"answer": "...",
"confidence": 0.88,
"task": "...",
"tools": ["..."],
"evidence": [{"label": "...", "url": "..."}],
"execution_trace": {}
}This lets the UI remain stable while models are upgraded.
User → Input Validator → Agent Controller → Task Router → Specialist Model Registry → Model Execution → Evidence Generator → Result Integrator → Answer + Evidence + Confidence + Trace.
"SatQuery AI is designed as an agentic orchestration layer rather than a single generic VLM. It validates the input configuration, classifies the user's intent, selects a specialist workflow, integrates textual and spatial evidence, and exposes an auditable execution trace."