In Akan folklore, Anansi the spider tricked the sky-god into giving him every story in the world, then wove them into a single web so people could share them. This Anansi does the same thing for mini-grid operators, sans trickery — weaving the scattered threads of daily work (meters, maps, tickets, field conversations, tribal knowledge in Telegram groups) into one chat thread your team and customers can actually talk to.
Built and run in production by NXT Grid, open-sourced for the wider energy-access community.
The admin app displays the public name Mini-Grids Assistant on the login screen and in UI copy as of PR #115 (2026-08-20); the codebase and repo remain named Anansi.
Everything below is one DigitalOcean App Platform app. This section is the map of what crosses its edges — what can call in, what it calls out to, and where state actually lives. Most of the subtle bugs in this repo have come from getting one of these boundaries wrong, so it's worth reading before changing anything that touches routing, auth, or persistence.
A single hostname fronts two services, split by path prefix. Rules are evaluated in order, and the last one is a catch-all, so a new path that doesn't match an earlier rule silently lands on the admin app (usually appearing as its login page rather than an error).
| Path prefix | Service | Prefix forwarded? | What it's for |
|---|---|---|---|
/chat |
chat-orchestrator | preserved | Telegram webhook + /chat/notify external alert passthrough |
/mini-app |
chat-orchestrator | preserved | Telegram Mini App UI |
/api/mini-app |
chat-orchestrator | preserved | Mini App backend calls |
/webhook |
chat-orchestrator | preserved | Inbound Jira webhooks |
/mcp-gateway |
chat-orchestrator | preserved | MCP endpoint + its OAuth routes (see MCP gateway below) |
/.well-known/oauth-*/mcp-gateway |
chat-orchestrator | preserved | RFC 8414/9728 OAuth discovery |
/ |
anansi-app | stripped | NiceGUI admin UI (catch-all) |
preserve_path_prefix is not cosmetic. DigitalOcean strips the matched prefix before forwarding unless it's set. Whether you want it depends entirely on whether the service's own routes include the prefix: chat-orchestrator serves real /chat/... routes, so it needs the prefix preserved. The MCP gateway's own routes are bare (/mcp, /healthz, /oauth/...), but it is mounted at /mcp-gateway inside chat-orchestrator, and Starlette's Mount strips the prefix itself — so ingress has to preserve it for the mount to match at all. Getting this backwards produces a 404 on every request, or a redirect to a URL missing the prefix entirely.
| Store | Access | Writeable? | Holds |
|---|---|---|---|
Auth DB (AUTH_DB_*) |
direct Postgres (asyncpg, pooler port) | read-only | public.accounts, grids, organizations, dcus — the operational/grid records that drive permissions |
Chat DB (CHAT_DB_URL + CHAT_DB_SERVICE_KEY, also CHAT_DB_POSTGRES_URL) |
Supabase / PostgREST | writeable | conversations, RAG chunks, prompts, skills, escalations, and anything this repo migrates |
Timescale (TIMESCALE_*) |
direct Postgres | reads | time-series telemetry |
DO Spaces (DO_SPACES_*) |
S3 API | writeable | object storage |
Two rules follow from this and are easy to get wrong:
db/migrations/targets the Chat DB only. Anything needing a new table must live there, because the Auth DB is read-only to this app — a feature that writes toAUTH_DB_*cannot work, no matter how the connection is configured.- Merging a migration does not apply it. Someone with Chat DB access runs it separately. An unapplied migration usually fails at runtime as
UndefinedTableError, not at deploy time.
Grouped by what they're for; each is gated by its own env vars and most by an *_ENABLED flag.
- LLMs — Gemini (
GEMINI_*,MODEL_THINKING/FAST/LITE), OpenRouter as an alternate provider, Langfuse for tracing - Google Workspace — OAuth sign-in (
AUTH_CLIENT_ID/SECRET), plus Docs/Drive/Sheets via a service account (GOOGLE_SERVICE_ACCOUNT_JSON,*_DOC_ID) and an Apps Script helper (ANANSI_HELPER_*) - Field/grid systems — VRM (
VRM_TOKEN, MQTT), ChirpStack, Calin v1/v2, metering and Tiamat APIs - Ops tooling — Jira (tickets + webhooks), Grafana (dashboards), Loki (logs), GitHub (
GITHUB_TOKEN, codebase search), Tavily (web search) - Messaging — Telegram bot API
- Payments — payment processor (
PAYMENT_PROCESSOR_*)
Each of these is reachable to the assistant as an MCP server under mcp_servers/servers/; the gateway's tiers.py decides which are exposed to external MCP clients and which never are.
Customers message the bot to check balance, buy tokens, report no-power, or get a token resent. Staff use the same bot to commission new meters, unassign them, change power limits, and resolve disputes. Anansi talks to your meter backend (Metering Platform by default, swappable via MCP) and applies the right approval rules depending on whether the requester is a customer or a staff member.
![]() Customer reports a meter issue |
![]() Issue resolved end-to-end |
The /lpp expert generates a Light Preliminary Package for a candidate site. Give it a GPS point and a site name; it pulls the community boundary from the GRID3 Nigeria settlement-extents dataset, runs the layout engine to place poles and lines, renders a map, and drafts a Google Doc package. The output is structured so the design can feed directly into a downstream Bill-of-Materials tool without re-keying.
![]() Site selection & community boundary |
![]() Generated distribution layout |
![]() Power demand heat map |
For the generation side, Anansi renders the power-plant footprint itself: PV array blocks with plinths, earth pits, lightning-arrester coverage circles, DC and AC cable runs with lengths, the Victron cabin, feeder pillar, VSAT, and the fenced site boundary with a gate. Module count, achieved vs target kWp, and cable lengths (including contingency) are summarised in the title block, so the same image doubles as a quick BoM sanity-check.
![]() PV arrays, earth pits, lightning coverage, cable trenches & cabin |
/analyze, /kpi, and /report let staff ask "how did Site X perform last week?" in plain English. Anansi pulls from TimescaleDB, the Victron VRM API (solar inverter telemetry), and your operational DB, then returns charts plus a written summary. Reports can be scheduled — e.g. every Monday at 9am to a specific Telegram group.
![]() Single-grid status report |
![]() Multi-grid KPI overview |
![]() Grid issue diagnosis |
Conversations that need a human are routed to the right internal Telegram group automatically, and tracked as a ticket with the full transcript attached — as a JIRA issue when Jira is configured and healthy, or in an internal ledger when it isn't, so escalations keep working (and stay recoverable) with or without Jira. The same pipeline ingests your historical Telegram support chats, Google Drive docs, and GitHub repos into a GraphRAG index — so the bot answers from your actual past decisions, not generic LLM knowledge.
Escalation tracking |
![]() Bot performance & escalation analytics |
Skills are multi-step automations that operators author by having a normal conversation with the bot through the admin app's builder UI, with no code or prompt engineering required. Each message becomes one step, and steps can be LLM instructions or pre-built function calls, so existing tool calls can be composed into a flow alongside prose instructions. Saved Skills can be scheduled (recurring or one-off) and run unattended end-to-end with no operator present to resume or confirm mid-run. Skills are the user-buildable counterpart to the hardcoded Expert Subagents, which handle workflows where step ordering is enforced in code.
Everything the bot is told directly — not generated by the model, not retrieved by RAG — lives on one admin page, grouped by where it comes from. Built-in modules (who's in a grid, the knowledge graph, episodic memory) are generated by the code per request; you choose which prompts use them, but can't edit their content. Curated modules are typed straight into the admin UI — this app is the source of truth. External modules attach a Google Doc or Sheet: the content stays in Drive, fetched fresh on every request and filtered to who's asking, so a procedures doc your ops team already maintains stays current with zero duplication. Every module is scoped (everywhere, or one organization) and, for an attached document, has an explicit audience — mirror the document's own sharing, or publish it to everyone the prompt serves — so a scoping mistake fails loudly instead of quietly leaking, or quietly omitting, content. A module isn't limited to prompts either: the same picker lives on a skill's own Context card in the Workflows editor, so a skill's steps can draw on the exact same curated knowledge a prompt does, resolved fresh at the start of each run. That includes generative prompts, not just conversational ones — pin a house-style guide to doc_editing.edit_highlighted and every @anansi-chatbot comment or chat instruction that edits a Google Doc inlines it. Scope still applies: a site- or organization-scoped module only reaches an edit whose caller resolved that grid or organization (an expert-workflow edit does; a raw chat edit today does not, so pin global modules there).
![]() Built-in, Curated, External — grouped by who's the source of truth |
![]() Attaching a Doc — live preview, access checked against your own Drive permissions |
Anansi is a provider-aware LLM chat orchestrator at its core — MCP tools, RAG, expert workflows, and a generative layer meant to be reshaped by the people running it, not just the people who built it. Three admin-app surfaces do that reshaping, live, with no redeploy and no code change: prompts (what the bot is instructed to do — your ops team edits any prompt in place; a bundled default always ships with the code, and a Google Doc can still be attached per prompt for teams that prefer editing there), context modules (what it's told as fact — see Context above; a module can attach to a prompt or to a skill), and Skills (multi-step automations it can run — see Skills above). That's the surface area ops and leadership actually touch to change what the bot does and knows; the sections below are for the engineers keeping the surface itself running. Gemini remains the default generation provider, and the shared LLM gateway can also be pointed at OpenRouter for OpenAI-style chat completions. The mini-grid focus comes from the tools and embellishments layered on top, and from the "messenger-first" assumption that field staff and customers live in chat apps, not dashboards. Telegram is the primary surface today; WhatsApp is on the roadmap but not yet supported.
Project structure:
chat_orchestrator/- Main chat orchestration service; Gemini is the default providermcp_servers/- MCP tool servers (grid design, meters, equipment control/diagnostics, JIRA, Grafana, payments, solar, knowledge, reference, and more — see MCP Servers below)rag_pipeline/- Knowledge ingestion from GitHub, Google Drive, Telegramshared/prompts/- The prompt library: bundled defaults, DB overrides, Google Doc attachments, and tagged knowledge modules composed into prompt contextshared/- Common utilities (auth, database, logging, Google Docs fetching, provider-neutral LLM gateways)anansi_app/- NiceGUI admin UI for chat history, broadcasts, grid design, Skills, and settingsmini_app/- Vite customer chat widget embedded in operator portalsllms.txt- Short repo map for LLM-assisted setup and onboarding. Keep it in sync when README setup, provider configuration, or major component paths change.
- Python 3.11+
- Google AI Studio or Gemini API key for the default provider (Get one)
- Optional OpenRouter API key if you want to exercise the shared OpenRouter generation gateway
- Google Cloud service account with Docs API enabled
- Supabase account
# Create service account and enable APIs
# See: https://console.cloud.google.com/iam-admin/serviceaccounts
# Enable these APIs:
# - Google Docs API
# - Google Drive API
# Download credentials JSONEvery prompt Anansi sends to a model — customer instructions, staff instructions, expert definitions, and about twenty smaller prompts used by individual workflows — lives in shared/prompts/library/*.prompt and ships with the code. A fresh clone works immediately with these generic defaults; nothing in this step is required to get started.
To customize for your organization, sign in to the admin app and open Prompts (/prompts). Pick an overridable prompt (most are; a handful of policy/routing prompts are locked and shipped-only by design — see the prompt's own detail view), edit its body, and either save a draft or publish it live. Changes take effect within about a minute, no redeploy.
If your team prefers editing in a Google Doc instead of the in-app editor, you can still attach one per prompt — the doc becomes that prompt's live source, with the bundled file as the fallback if the doc becomes unreachable. This is optional; the in-app Prompts page is now the primary way to edit.
If you do attach a Google Doc, give it this structure so section parsing works:
[Optional title page]
[PAGE BREAK]
Heading 1: System Instructions
You are a helpful assistant for [Company Name].
Be professional, empathetic, and accurate.
Heading 1: QnA Knowledge Base
Q: What are your business hours?
A: Monday-Friday, 9 AM - 5 PM EST.
Heading 1: Example Conversations
User: I can't log in
Assistant: I understand that's frustrating. Let me help...
Share the doc with your service account:
- Get the service account email from the credentials JSON:
"client_email": "..." - Share the doc with Viewer access
# Clone and setup shared utilities
git clone <repository-url>
cd anansi
./setup_shared.sh
# Chat Orchestrator
cd chat_orchestrator
python3 -m venv .venv
source .venv/bin/activate
pip install -e .[dev]
# Install pre-commit hooks (code quality checks on every commit)
pre-commit install
# Configure environment
cp .env.example .envEdit .env:
# Required
LLM_PROVIDER=gemini
GOOGLE_API_KEY=your-gemini-api-key
GOOGLE_SERVICE_ACCOUNT_JSON='{"type":"service_account",...}' # only needed if you attach Google Docs to prompts
# Optional — only if attaching a Google Doc to customer.system / staff.system
# instead of editing them from the Prompts admin page (see step 2 above)
CUSTOMER_SUPPORT_DOC_ID=1abc123xyz456 # From GDoc URL
STAFF_SUPPORT_DOC_ID=1def789uvw012 # From GDoc URL
# Who may edit/publish prompts from the Prompts admin page (comma-separated
# emails). All three are optional; unset means nobody can edit anything —
# the bot still runs fine on bundled defaults.
PROMPT_EDITORS_OPS=ops@example.com
PROMPT_EDITORS_ENG=eng@example.com
PROMPT_ADMINS=admin@example.com
# Chat Database (for conversations and RAG)
CHAT_DB_URL=https://your-project.supabase.co
CHAT_DB_SERVICE_KEY=your-service-role-key
# Auth Database (read-only for user lookup)
AUTH_DB_DIRECT_CONNECTION=true
AUTH_DB_HOST=db.your-auth-project.supabase.co
AUTH_DB_PORT=6543
AUTH_DB_NAME=postgres
AUTH_DB_USER=readonly_user
AUTH_DB_PASSWORD=your_password
AUTH_DB_SSL_MODE=require
# Gemini Config
# MODEL_THINKING/MODEL_FAST/MODEL_LITE are required -- each prompt declares
# which of the 3 tiers it uses (shared/llm/model_tiers.py), and resolving an
# unset tier raises rather than silently falling back.
MODEL_THINKING=gemini-pro-latest
MODEL_FAST=gemini-flash-latest
MODEL_LITE=gemini-2.5-flash-lite
FALLBACK_MODEL=gemini-2.5-flash-lite
GEMINI_TEMPERATURE=0.7
# OpenRouter compatibility (optional; keep LLM_PROVIDER=gemini unless testing it)
# LLM_PROVIDER=openrouter
# OPENROUTER_API_KEY=your-openrouter-api-key
# OPEN_ROUTER_BEARER_TOKEN is also accepted as a local alias
# OPENROUTER_MODEL=google/gemini-2.5-flash
# OPENROUTER_PROVIDER_ORDER=google-vertex
# OPENROUTER_ALLOW_FALLBACKS=false
# OPENROUTER_HTTP_REFERER=https://yourapp.example.com
# OPENROUTER_APP_TITLE=Anansi# Development (all services)
./dev.sh
# Or orchestrator only
cd chat_orchestrator && source .venv/bin/activate
uvicorn orchestrator.api.app:app --host 0.0.0.0 --port 8000 --reload
# Test endpoint
curl http://localhost:8000/healthAfter the basic setup works, these are the things every operator should configure before going live:
STAFF_ORG_ID controls which organization_id in your accounts table gets staff-mode access (full tools, staff instructions). The default is 2 — change it to match your own database:
STAFF_ORG_ID=5 # whatever your staff org's ID is in the accounts tableStaff users see the STAFF_SUPPORT_DOC_ID instructions and have access to all MCP tools. Everyone else gets CUSTOMER_SUPPORT_DOC_ID and a limited tool set.
Set your Telegram bot's @handle so group-chat mention detection works correctly:
TELEGRAM_BOT_USERNAME=YourBotName # without the @ prefixWithout this, the bot won't respond when mentioned by name in group chats.
The bundled prompts under shared/prompts/library/*.prompt (including customer.system and staff.system) are intentionally generic and will not reflect your organization's actual support process. For a real deployment, edit them from the Prompts admin page (/prompts) — no env vars required.
If you'd rather keep editing in Google Docs, the legacy doc-id env vars still work exactly as before and take precedence over the bundled default (though a DB override made from the Prompts page takes precedence over the doc, if you use both):
CUSTOMER_SUPPORT_DOC_ID=<your-customer-doc-id>
STAFF_SUPPORT_DOC_ID=<your-staff-doc-id>
EXPERT_INSTRUCTIONS_DOC_ID=<your-expert-doc-id>Gemini is the supported default:
LLM_PROVIDER=gemini
GOOGLE_API_KEY=your-gemini-api-key
MODEL_THINKING=gemini-pro-latest
MODEL_FAST=gemini-flash-latest
MODEL_LITE=gemini-2.5-flash-liteThe shared generation gateway can also call OpenRouter using OpenAI-compatible chat completions:
LLM_PROVIDER=openrouter
OPENROUTER_API_KEY=your-openrouter-api-key
# OPEN_ROUTER_BEARER_TOKEN is also accepted as a local alias
OPENROUTER_MODEL=google/gemini-2.5-flash
OPENROUTER_PROVIDER_ORDER=google-vertex
OPENROUTER_ALLOW_FALLBACKS=false
OPENROUTER_HTTP_REFERER=https://yourapp.example.com
OPENROUTER_APP_TITLE=AnansiKeep LLM_PROVIDER=gemini for normal deployments until you intentionally test OpenRouter-backed generation paths. For OpenRouter BYOK/BYOL with Google Vertex, set OPENROUTER_PROVIDER_ORDER=google-vertex and OPENROUTER_ALLOW_FALLBACKS=false so requests do not fall back to other OpenRouter endpoints. The settings page discovers provider routes from the selected OpenRouter model using the normal OpenRouter access key. Gemini-specific orchestrator code remains available as the default backup path.
The AI Models & Providers settings page has a separate Use Jev for decisions switch and Jev decision model dropdown. The switch is off by default. It applies to customer-response verification (when VERIFICATION_ENABLED is on), issue-type classification, ticket-comment significance, multi-active-thread assignment (when THREAD_DISENTANGLEMENT_ENABLED is on), and conversation context filtering (when CONTEXT_FILTER_ENABLED is on). Gemini can remain the generation provider while these decisions use an OpenRouter key.
JEV_DECISIONS_ENABLED=false
JEV_DECISIONS_MODEL=~typesafe/jev-latest
OPENROUTER_API_KEY=your-openrouter-api-keyJev uses OpenRouter's POST /api/alpha/decisions, separate from chat completions. The configured model must be in the Jev family. On missing credit, a transport error, an invalid answer, or an uncertain decision, the existing LLM call remains the fallback. The response verifier uses Jev only to fast-pass clear messages; possible failures still go to the existing judge for written repair feedback. Long inputs and tool activity with no source evidence also use the existing judge. The technical-response sanitizer still uses a generative model. The OpenRouter provider-order controls above are for chat completions and are not sent to Decisions.
Before enabling the switch in production, fund the OpenRouter account and smoke-test one Choice, one Noul, and one batched request with the selected model. Compare the five adapters on labeled examples in the supported customer languages and calibrate their thresholds. The implementation's offline tests check the documented wire shape and fallback behavior, but no funded live Jev call has been run for this integration. Set JEV_DECISIONS_ENABLED=false to return every path to its previous behavior.
The shared/auth code references a column named is_generation_managed_by_nxt_grid in the grids table (via the MANAGED_GENERATION_COLUMN env var, defaulting to that name). This is an operator-specific column from the reference deployment. If your schema uses a different name (or doesn't have this concept), set MANAGED_GENERATION_COLUMN or update the default in shared/auth/auth_service.py.
Most MCP servers are disabled by default. Enable only what you have credentials for via the {SERVER_NAME}_ENABLED env vars. See mcp_servers/.env.example for the full list with documentation.
Every prompt resolves through shared.prompts.PROMPTS, in this order — each layer optional, falling through to the next on any failure:
1. DB override (Prompts admin page) — only if the prompt is overridable
↓ (if none, or lookup fails)
2. Attached Google Doc — only if a doc id is configured for this prompt
↓ (if none, or fetch fails)
3. Bundled default (shared/prompts/library/<id>.prompt) — always present
For a Google Doc source, steps in between:
a. Fetch via Docs API
b. Convert to Markdown (Heading 1-6 → # to ######, Bold → **text**, Italic → *text*)
c. Auto-strip title page, headers/footers, inline images
Whatever body wins resolution is then:
4. Parsed into sections by Heading 1 (per the prompt's declared `sections`)
5. Split: the named system-instructions section → systemInstruction field;
everything else (QnA, Examples, and any composed knowledge modules) →
first user message
6. Sent to Gemini API (default provider)
{
"systemInstruction": {
"parts": [{"text": "System Instructions section"}]
},
"contents": [
{
"role": "user",
"parts": [{"text": "QnA + Examples + Technical Knowledge sections"}]
},
{
"role": "user",
"parts": [{"text": "Actual user question"}]
}
]
}
Every render carries provenance (which prompt id, source, version, checksum produced it) into logs and, when LANGFUSE_ENABLED=true, the trace.
Determined by user's organization_id:
-
organization_id = STAFF_ORG_ID(env var, default2) → Staff mode (internal users)- Resolves the
staff.systemprompt - Access to all MCP tools
- Full system capabilities
- Resolves the
-
All other
organization_idvalues → Customer mode (external users)- Resolves the
customer.systemprompt - Limited to customer support tools
- Safe, scoped responses
- Resolves the
Both modes use identical processing pipeline.
sequenceDiagram
participant U as User (Telegram / API)
participant O as chat_orchestrator<br/>(FastAPI + LangGraph)
participant G as Gemini API<br/>(default provider)
participant M as mcp_servers<br/>(tool handlers)
participant DB as Databases<br/>(Supabase / TimescaleDB)
participant GD as Google Docs<br/>(system instructions)
U->>O: POST /chat {message}
O->>GD: Fetch system instructions doc
GD-->>O: Markdown sections
O->>DB: Load conversation history
O->>G: Chat request (systemInstruction + history + message)
G-->>O: Response or tool_call
alt tool call requested
O->>M: Execute tool (e.g. get_grid_status)
M->>DB: Query data
DB-->>M: Results
M-->>O: Tool result
O->>G: Continue with tool result
G-->>O: Final response
end
O->>DB: Persist message
O-->>U: Response
Telegram / Web Client
│
▼
┌───────────────────────┐
│ chat_orchestrator │ FastAPI — main chat orchestration
│ (port 8000) │ LangGraph stategraph, expert workflows,
│ │ MCP tool execution
└───────┬───────────────┘
│ Python imports (monorepo)
▼
┌───────────────────────┐
│ mcp_servers │ MCP tool servers — JIRA, Grafana,
│ │ customer data, equipment control,
│ │ meters, schedule, payments, knowledge
└───────────────────────┘
┌───────────────────────┐
│ rag_pipeline │ Document ingestion — GitHub, Google
│ │ Drive, Telegram; GraphRAG embeddings
└───────────────────────┘
┌───────────────────────┐
│ anansi_app │ NiceGUI admin UI — chat history, grid
│ (port 8501) │ design, Skills, settings, scheduler
└───────────────────────┘
┌───────────────────────┐
│ mini_app │ Vite/Vanilla JS customer chat widget
│ (served via bot) │ (embedded in operator portals)
└───────────────────────┘
Purpose: Orchestrate LLM conversations with dynamic instructions; Gemini is the default provider
Key Features:
- Google Docs integration (single source of truth)
- Section parsing and markdown conversion
- Proper Gemini API usage for the default path (
systemInstructionfield) - Context injection as first user message
- Multi-turn conversation loops
- Parallel tool execution
Location: chat_orchestrator/
Purpose: Ingest and index knowledge from multiple sources
Two Ingestion Methods:
-
Batch Indexers (CLI) - Bulk ingestion for initial setup
- GitHub repositories (code + docs)
- Google Drive folders (docs, PDFs, spreadsheets)
- Telegram chats (messages, topics)
-
Interactive Ingestion (
/learn_ragcommand) - Individual documents via Telegram- Google Docs → Markdown (preserves formatting)
- PDFs →
pymupdf4llm(markdown output, tables) - LLM-based document classification
- Procedure matching for support examples
- User approval before storage
GraphRAG Features:
- Semantic chunking (~512 tokens)
- Vector embeddings (768-dim, Google AI Studio)
- Hybrid retrieval (vector similarity + full-text search)
- Entity extraction
- Relationship mapping
- Agentic graph query tools for entity-relationship exploration
- Procedure tagging for filtered retrieval
Location: rag_pipeline/ (batch), chat_orchestrator/orchestrator/experts/handlers/ingestion_expert/ (interactive)
Purpose: Tool integration via Model Context Protocol
Available Tools:
- Customer - Customer-facing tools for payment and commissioning status
- Meters - Meter management and operations
- Equipment Control - Equipment control operations
- Equipment Diagnostics - Production equipment diagnostics, historical analysis, charts, and monitoring
- Grid Design - Grid design and Bill of Materials generation
- Solar - Solar potential assessment using the Global Solar Atlas API
- JIRA - Jira analysis and comment processing
- Grafana - Grafana dashboard panel rendering
- Payment Processor - Payment processor transaction status checks
- Knowledge - Knowledge base summarization and exploration tools
- Reference - Nigerian import tariff, prohibition list, and standards lookups (staff only)
- Schedule - Command scheduling (staff only)
- Meta - Bot performance analytics (staff only)
Location: mcp_servers/
Purpose: Handle complex, multi-step workflows with structured state management
How It Works:
- Triggered by slash commands (
/lpp,/analyze,/kpi) - Maintain workflow state in work packets (database-persisted)
- Can pause for user input and resume later
- Failed workflows can be retried or abandoned
Available Experts:
| Command | Expert | Description |
|---|---|---|
/lpp |
Package Generator | Generate Light Preliminary Packages |
/analyze |
Grid Analyst | Analyze grid performance and faults |
/kpi, /report |
Grid Analyst | Generate KPI reports |
/csize |
Community Sizing | Detect community at GPS coordinates and estimate solar sizing |
/sign |
Signing | Request a signature on a Drive PDF from a named person |
/gtr |
Grids Technical Reviewer | Generate monthly technical review for grid(s) |
/codebase, /anansi |
Code Investigation | Investigate the platform or Anansi codebase for an issue |
/learn |
Context Ingestion | Teach the bot a fact it should always know (context module) |
/learn_rag |
Ingestion | Add a source document to the searchable knowledge base |
Expert Definition (the experts.definitions prompt — bundled by default, editable from the Prompts admin page, or a Google Doc via EXPERT_INSTRUCTIONS_DOC_ID):
# Expert: package_generator
## Model
gemini-3-flash
## System Instructions
You are a specialist in creating Light Preliminary Packages...
## Tools
- google_docs
## Packet Types
- light_preliminary_package
## Packet: light_preliminary_package
### Workflow
[llm] parse_request - Extract site name
[function:generate_lpp_map] - Generate map
[function:copy_lpp_template] - Create document
[llm] summarize_result - Report to userPer-Expert Model Override: Add ## Model section to use a different Gemini model for that expert (e.g., gemini-3-flash).
Location: chat_orchestrator/orchestrator/experts/
Purpose: Operator-authored, code-free multi-step automations that run unattended on a schedule or on demand.
How It Works:
- Authored in the admin app's Skills builder: one user message in a conversation with the bot equals one step
- Steps are either LLM steps (a plain-language instruction) or function steps (a pre-built function call); authors can compose existing tool calls into a flow, not only prose instructions
- Saved Skills can be scheduled (recurring or one-off) or triggered like any other command, then run end-to-end unattended with no operator present to confirm or resume mid-run
- Runtime is
skill_runner.py, backed byskill_step_bindings.py,skill_schedule_dispatch.py,skill_summary.py, andskill_validation.py - Several original Expert Subagents (above) with no real step logic are being migrated into Skills; four genuine multi-step Experts (
context_expert,grids_technical_reviewer,ingestion_expert,package_generator) stay as hardcoded Experts
Location: anansi_app/nicegui_app/pages/skills.py, skill_builder.py (admin app); chat_orchestrator/orchestrator/experts/skill_runner.py and supporting files; DB migrations 0011_skills.sql, 0013_skill_scheduling.sql, 0025_skill_draft_status.sql, 0026_user_schedules_skill_unique.sql
User Request
↓
Instructions Provider
├─ Resolve via the prompt library (DB override → Google Doc → bundled)
├─ Parse sections
└─ Split: System Instructions vs Context
↓
Conversation Orchestrator
├─ systemInstruction field
├─ Context as first user message
└─ User request as last message
↓
Gemini API (default provider)
├─ Calls tools if needed (MCP servers)
└─ Retrieves RAG context if needed
↓
Response
| File | Purpose |
|---|---|
orchestrator/services/conversation.py |
Main conversation orchestration |
orchestrator/graphs/conversation_graph.py |
LangGraph stategraph |
shared/prompts/ |
The prompt library: bundled defaults, DB overrides, Doc attachments, knowledge modules |
orchestrator/services/instructions_provider.py |
Composes customer/staff prompts into system instructions + context |
orchestrator/services/tool_executor.py |
MCP tool execution |
orchestrator/services/command_registry.py |
Slash command definitions |
mcp_servers/tool_definitions.json |
All tool definitions (source of truth) |
mcp_servers/server_registry.py |
MCP server registry |
shared/auth/auth_service.py |
Authentication |
Two kinds of configuration exist. Credentials and connection strings (API
keys, database URLs, OAuth secrets) come from the host environment — set them
in your platform's env var UI or a local .env and they are never written
back. Operator-tunable flags (feature toggles, model choices, timeouts,
layout parameters) live in shared/config/flag_registry.py,
are documented in the generated shared/config/flags.env.example,
and are normally set through the anansi_app Settings page rather than by
editing the environment directly.
The settings page's own Deployment Readiness panel reports exactly what a given environment is still missing, grouped by capability rather than by variable name. The tiers below are what it checks, in the order a new deployment typically reaches them:
GRID_DESIGN_DEV_NO_AUTH=1 # bypasses Google OAuth entirely — never set this in productionNothing else is required to boot anansi_app and reach /settings locally.
GOOGLE_CLIENT_ID=your-oauth-client-id.apps.googleusercontent.com
GOOGLE_CLIENT_SECRET=your-oauth-client-secret
AUTH_REDIRECT_URI=http://localhost:8501/oauth2callback
ALLOWED_VIEWER_EMAILS=admin@example.comWithout a configured remote backend, the Settings page is read-only. This is deliberate: a file written inside the admin container cannot update the bot's separate process or survive replacement on most container hosts.
For local settings-page development only, opt into the env-file backend:
SETTINGS_BACKEND=envfile
SETTINGS_FILE=.env.settingsThis records changes in the selected file but does not load them into other services automatically. Set their environment and restart them explicitly.
For DigitalOcean, configure the live app-spec backend (redeploys on save):
DIGITALOCEAN_APP_ID=your-do-app-id
DIGITALOCEAN_API_TOKEN=your-do-api-tokenGOOGLE_API_KEY=your-gemini-api-key
TELEGRAM_BOT_TOKEN=your-telegram-bot-token
CHAT_DB_URL=https://your-project.supabase.co # or SUPABASE_URL
CHAT_DB_SERVICE_KEY=your-service-role-key # or SUPABASE_KEY
API_KEY=your-orchestrator-api-key
SESSION_ID_SECRET=generate-a-random-secret
# Authentication — Option A: direct PostgreSQL (recommended)
AUTH_DB_DIRECT_CONNECTION=true
AUTH_DB_HOST=db.your-auth-project.supabase.co
AUTH_DB_PORT=6543
AUTH_DB_NAME=postgres
AUTH_DB_USER=readonly_user
AUTH_DB_PASSWORD=your_password
AUTH_DB_SSL_MODE=require
# Option B: Supabase client
# AUTH_SUPABASE_URL=https://your-auth-project.supabase.co
# AUTH_SUPABASE_KEY=your_auth_service_key
# System instructions
GOOGLE_SERVICE_ACCOUNT_JSON='{"type":"service_account",...}'
CUSTOMER_SUPPORT_DOC_ID=1abc123xyz456
STAFF_SUPPORT_DOC_ID=1def789uvw012
EXPERT_INSTRUCTIONS_DOC_ID=1ghi456jkl789 # Expert definitions# Jira (escalations; without these, escalations still post to Telegram and
# are tracked in the internal ticket ledger)
JIRA_BASE_URL=https://your-domain.atlassian.net
JIRA_USERNAME=your-email@example.com
JIRA_API_TOKEN=your-api-token
JIRA_WEBHOOK_SECRET=a-long-random-string # see "Jira Webhook" below
# Grafana (dashboard/panel tools)
GRAFANA_URL=http://localhost:3000
GRAFANA_USERNAME=admin
GRAFANA_PASSWORD=your-grafana-password
# POST /chat/notify (Grafana / n8n / VRM passthrough)
NOTIFY_SHARED_SECRET=your-shared-secret
# RAG pipeline
GITHUB_TOKEN=ghp_your-token
GITHUB_REPO=owner/repo
TELEGRAM_API_ID=12345678
TELEGRAM_API_HASH=abc123...Every other tunable (which MCP servers are enabled, model choices, layout
geometry, RAG toggles, and the read/write gate per MCP server) has a sensible
default and is listed in full, with descriptions, in
shared/config/flags.env.example.
Escalations always post to the internal Telegram support group; whether they're also tracked as a Jira ticket or an internal ledger entry is decided per-call by TicketService, independent of Jira being configured at all. All of these are optional — sensible defaults apply if unset — and are managed like any other operator flag (see shared/config/flag_registry.py / shared/config/flags.env.example for the full, generated list):
# 'auto' (default): Jira if JIRA_* creds are set and Jira answers a health probe,
# else the internal ledger (internal_tickets / internal_ticket_comments in chat_db).
# 'jira': Jira if creds are present, else internal (never hard-fails).
# 'internal': always internal — an ops kill-switch, e.g. during a Jira outage.
TICKET_BACKEND_OVERRIDE=auto
# Backend for tickets filed via POST /chat/notify (see below). Independent of
# TICKET_BACKEND_OVERRIDE: defaults to 'internal' so Grafana/n8n/VRM alerts
# never land in the Jira project unless you opt into 'auto'.
NOTIFY_TICKETS_BACKEND=internal
# Prefix for internal ticket refs, e.g. 'TKT' -> 'TKT-000123'.
INTERNAL_TICKET_PREFIX=TKT
# How long the Jira health probe result is cached before re-checking (seconds).
JIRA_HEALTHCHECK_TTL_SECONDS=60Without any Jira credentials configured, escalations, the staff Track/Close buttons, and the daily sweep all work exactly the same, filing and updating internal tickets instead. The on-call schedule (get_on_call/add_on_call_override, JSM Ops) stays Jira-dependent by design — with Jira offline, on-call queries return a clean "unavailable" response instead of an error.
Ticket references are backend-neutral after creation: TKT-* tickets are read, commented on, and closed through the same ticket tools even if NOTIFY_TICKETS_BACKEND is later changed to Jira; Jira references remain Jira-backed. Assignment and arbitrary workflow transitions are Jira-only.
POST /chat/notify ticketing: the existing alert-forwarding endpoint (source, grid_name, text, ...) accepts a few more optional fields:
ticket_id— omit for today's plain passthrough (unchanged). Pass""to file a new ticket from this notification (response includesticket_ref). Pass"auto"to let Anansi decide whether this alert is new, relates to an already-open ticket on this grid, or is an exact re-fire — see "Alert correlation" below. Pass an existing ref (e.g."TKT-000123"or"OPS-55") to append the notification as a comment on that ticket; an unresolvable ref returns404.close— with a populatedticket_id, also transition that ticket to done. Ignored forticket_id="auto".alert— optional structured facts (subject,alert_type,details,severity,component_kind/component_key/component_label,fired_at,rule_id) used byticket_id="auto"correlation. Every field is independently derivable fromtext/subjectwhen omitted; pass what you already have (e.g. n8n's extracted MPPT/DCU id) and Anansi fills in the rest.
The alert is still forwarded to Telegram in every case; ticket_id/close/alert only control the ticketing side effect.
One root cause (e.g. a grid stuck OFF/Unknown for hours) can otherwise produce a storm of separate tickets — every dependent MPPT/DCU alert filing its own issue. ticket_id="auto" groups an incoming alert against a grid's already-open tickets instead: an LLM (given the grid's deterministic operational facts and the candidate open tickets) decides new / amend (a different affected component of the same issue — appended to the existing ticket's affected-components list) / duplicate (the exact same component re-firing — silent, occurrence-counted only). See docs/superpowers/plans/2026-07-27-smart-alert-correlation-notify.md for the full design.
The response for ticket_id="auto" adds decision ("new"|"amend"|"duplicate"), correlated_with (the ticket this alert was matched against, or null for a new ticket), confidence, and decided_by ("replay"|"flag_off"|"no_candidates"|"signature"|"llm"|"fallback") alongside ticket_ref.
Fail-open guarantee: every failure mode — the LLM timing out or erroring, an unparseable response, a correlation-store outage, a per-grid lock timeout — falls back to filing a plain new ticket (decided_by="fallback"), the same as ticket_id="". Correlation only ever adds grouping on top; it can never cause an alert to be dropped.
# Choose the Jira project used when the Jira backend is selected.
JIRA_PROJECT_KEY=OPS
# Keep /chat/notify alerts internal by default; select auto to use healthy Jira.
NOTIFY_TICKETS_BACKEND=internal
# Leave correlation enabled. Set false only to bypass it and file a plain ticket.
ALERT_CORRELATION_ENABLED=trueAlert setup: set JIRA_PROJECT_KEY, choose NOTIFY_TICKETS_BACKEND, and leave ALERT_CORRELATION_ENABLED enabled unless you intentionally need the plain-ticket bypass. When Jira is selected but its project has no compatible issue type, /notify fails open to an internal TKT-* ticket.
Concurrency caveat: correlation serializes decisions per grid with an in-process asyncio.Lock, which is correct for the current single-process deployment (chat_orchestrator/Dockerfile runs uvicorn with no --workers). At instance_count > 1 (or with --workers), this stops serializing across processes — several alerts for the same grid arriving at once across instances can still each see "no open candidate" and file separate tickets. A distributed lease table is the documented follow-up (see the plan's "Concurrency" section); don't scale this endpoint horizontally without addressing it first.
Operational runbook:
- Disable correlation (keep ticketing):
ALERT_CORRELATION_ENABLED=false— everyticket_id="auto"request still files a plain ticket. - Disable ticketing into Jira (keep correlation):
NOTIFY_TICKETS_BACKEND=internal— alert tickets stay ininternal_tickets/ticket_correlations, never reach the Jira project. - Disable the endpoint entirely:
NOTIFY_ENDPOINT_ENABLED=false—/notify503s. - Inspect a decision: the
ticket_correlation_eventstable (or the ticket's "Decision history" section on its admin Tickets page detail view) has every decision,decided_by, confidence, and reason for a ticket_ref.
From Google Doc URL:
https://docs.google.com/document/d/1abc123xyz456/edit
^^^^^^^^^^^^^^
This is the doc ID
The Auth DB is your existing business database (users, organizations, sites). The bot connects read-only via asyncpg.
Minimum required tables (in public schema):
| Table | Columns used |
|---|---|
accounts |
id, email, telegram_id, organization_id, deleted_at |
organizations |
id, name, developer_group_telegram_chat_id |
grids |
id, name, organization_id, internal_telegram_group_chat_id, internal_telegram_group_thread_id, deleted_at |
dcus |
id, grid_id, deleted_at |
Create a read-only user:
CREATE USER anansi_readonly WITH PASSWORD 'your-strong-password';
GRANT CONNECT ON DATABASE postgres TO anansi_readonly;
GRANT USAGE ON SCHEMA public TO anansi_readonly;
GRANT SELECT ON public.accounts TO anansi_readonly;
GRANT SELECT ON public.organizations TO anansi_readonly;
GRANT SELECT ON public.grids TO anansi_readonly;
GRANT SELECT ON public.dcus TO anansi_readonly;
-- Add more grants as you enable optional MCP tool serversIf using Supabase for the Auth DB, use port 6543 (PgBouncer) and set statement_cache_size=0. Do not use make_readonly or the PostgREST client for this database.
| Table | Columns used | Purpose |
|---|---|---|
accounts |
id, email, telegram_id, organization_id, deleted_at |
Map Telegram user → org |
organizations |
id, name, developer_group_telegram_chat_id |
Org lookup and staff chat ID |
grids |
id, name, organization_id, internal_telegram_group_chat_id, internal_telegram_group_thread_id, deleted_at |
Map grid → Telegram group |
dcus |
id, grid_id, deleted_at |
Device lookup for grid tools |
Anansi only reads these tables (never writes). Add GRANT SELECT ON ... for any additional tables your MCP servers query.
- Create a Supabase project at supabase.com.
- Enable
pgvector:CREATE EXTENSION IF NOT EXISTS "vector"; - Run the bootstrap schema in Supabase SQL Editor:
# Copy and execute db/schema/chat_db.sql in the Supabase SQL Editor # (Project → SQL Editor → New query → paste → Run)
- From Project Settings → API, copy the project URL and
service_rolekey (not theanonkey).
| Table | Purpose |
|---|---|
chat_sessions |
One row per conversation thread (keyed by Telegram chat/thread ID or API session) |
chat_messages |
Full message history with role, content, and tool call records |
agent_work_packets |
State for multi-step expert workflows (paused, running, completed, failed) |
agent_work_packet_logs |
Execution log per workflow run — step timings, success/failure |
pending_decisions |
Multi-turn decision state (e.g. "duplicate detected — resume or start fresh?") |
documents |
RAG document store — metadata, access control, embeddings |
document_chunks |
Chunked text with vector embeddings (pgvector) |
The full schema is in db/schema/chat_db.sql and can be applied in one step via the Supabase SQL Editor.
The RAG schema (documents, chunks, vector index) is included in the main Chat DB schema.
Run db/schema/chat_db.sql in Supabase SQL Editor if you haven't already — no separate RAG migration needed.
cd rag_pipeline
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
# Configure
cp .env.example .env
# Add GOOGLE_API_KEY, SUPABASE credentials, etc.GitHub:
python ingestion/github_indexer_v2.py \
--repo owner/repo \
--source-name "Company Codebase"Google Drive:
python ingestion/gdrive_indexer_v2.py \
--folder-id YOUR_FOLDER_ID \
--source-name "Company Docs"Telegram:
python ingestion/telegram_indexer_v2.py \
--folder-id TELEGRAM_EXPORTS_FOLDER_ID \
--source-name "Support Chats"RAG is automatically used by the chat orchestrator when rag.enabled=true in settings.
One app, two services (see .do/app.example.yaml; the live reference deployment's chat service is named anansi-bot — adjust to match your own app spec):
| Component | Type | Description |
|---|---|---|
| chat-orchestrator | Service | Chat orchestration + MCP tools (consolidated; Gemini default). Also hosts the MCP gateway, mounted at /mcp-gateway — not a separate service. |
| anansi-app | Service | NiceGUI admin UI — chat history, grid design, Skills, settings. Also runs the broadcast-scheduler and Grafana-indexer daemons in-process (anansi_app/start.sh) — neither is a separate DO Job. |
# Deploy via doctl (SAFE pattern — never update directly from .do/app.yaml which has placeholders)
doctl apps spec get <app-id> > /tmp/live-spec.yaml
# edit /tmp/live-spec.yaml, then:
doctl apps update <app-id> --spec /tmp/live-spec.yaml
doctl apps create-deployment <app-id>
# Or push to main branch (auto-deploys)
git push origin mainExposes the same tools the assistant uses to external MCP clients (Claude Code, Claude Desktop, claude.ai, ChatGPT, Codex), scoped to whoever signed in — see Connecting a client below for the exact steps per client. A user connects once, for example:
claude mcp add --transport http anansi-mcp https://your-app.example.com/mcp-gateway/mcpand authenticates with their own Google work account in their own browser. No token is ever pasted anywhere, and no shared API key exists for this route.
Where it runs. Inside chat-orchestrator, not as its own service. That image already carries mcp_servers/ and already calls server_registry directly for every tool call, so a separate service duplicated the entire runtime — plus its own copy of the MCP server code, free to drift from the copy actually executing tools — to serve a few hundred requests a day. orchestrator/api/app.py mounts it at the path component of MCP_GATEWAY_BASE_URL, and starts its session manager from the app's own startup event (a mounted ASGI app never receives the lifespan scope, so this is not optional). Both are guarded: if the gateway cannot be built or started, it is skipped with a logged traceback and chat traffic is unaffected.
Turning it off. Unset either MCP_GATEWAY_BASE_URL or MCP_GATEWAY_TOKEN_SECRET. The gateway is then never mounted — no code change, no separate deploy, and nothing else in the service is affected.
Routes (all relative to /mcp-gateway, which the mount strips before the gateway sees it):
| Route | Purpose |
|---|---|
/mcp |
The MCP endpoint itself (Streamable HTTP). Requires a bearer token; without one it returns 401 + WWW-Authenticate, which is what triggers a client's OAuth flow |
/healthz |
Unauthenticated health check — /mcp-gateway/healthz publicly, and the fastest way to tell whether the mount succeeded |
/.well-known/oauth-protected-resource |
RFC 9728 — points clients at the authorization server |
/.well-known/oauth-authorization-server |
RFC 8414 — advertises the endpoints below |
RFC 8414 puts an issuer's metadata above its own path — an issuer of https://host/mcp-gateway publishes at https://host/.well-known/oauth-authorization-server/mcp-gateway, which no mount under /mcp-gateway could ever serve. Those two paths are registered on chat-orchestrator's root router instead (well_known_routes()), which is why they get their own ingress rules.
| /oauth/register | RFC 7591 dynamic client registration (Claude Code requires this and cannot be given a client ID by hand) |
| /oauth/authorize | Starts the Google leg. redirect_uri must be a loopback address (RFC 8252) or an exact match in MCP_GATEWAY_REDIRECT_ALLOWLIST |
| /oauth/google-callback | Google's redirect target — the one URI registered in Google Cloud Console |
| /oauth/token | PKCE-verified, single-use code exchange |
Domain context. Anansi's semantic value is not in its tools — it's in the system prompt and knowledge modules composed into every chat turn. A client calling the same tools without those gets the mechanics and none of the meaning. So the gateway hands the caller's own rendered prompt (staff.system or customer.system, chosen by the signed-in identity, with its attached knowledge modules composed in under the same RequestScope filtering the bot uses) to the client through InitializeResult.instructions — pushed in the initialize handshake, not something the client has to ask for.
Those prompts are written for the Telegram assistant, so they're prefixed with a preamble that frames them as a reference document and names the categories to disregard (channel formatting, inline buttons, escalation, orchestrator-internal tool names, chat turn-taking, media handling, session machinery), and closed with a one-line reminder. Sections are deliberately not marked for exclusion in the prompt bodies: those are co-edited by operators through the admin app, and per-section markers rot the first time someone rewrites a section without knowing they matter.
Because instructions is delivered once at initialize and MCP has no "instructions changed" notification, an edited prompt or knowledge module reaches an already-connected client only when it starts a new session. Tool results are unaffected — those are fetched fresh every call.
How authorization works. Two OAuth hops that are easy to conflate: the client ↔ this gateway (dynamic loopback redirect, PKCE, no Google involvement) and this gateway ↔ Google (one stable, pre-registered callback). The issued client_id is not a credential and is not persisted — for a public client the real gates are PKCE, loopback-only redirect_uri validation, the Google-verified identity, and the same grid_app.lib.perms whitelist the admin app uses.
Once connected, every request re-resolves the caller's organization and permissions from the database — nothing is cached for the life of a connection, so revoking someone takes effect on their next call. Tools are filtered per user by tiers.py (equipment control, payments and messaging are never exposed) and every scope-bearing argument is overwritten server-side rather than trusted from the caller.
Hosted clients (Claude Desktop, claude.ai, ChatGPT). A CLI client (Claude Code, Codex) redirects to a loopback address it controls, which needs no configuration. A hosted client has no local port to redirect to — its OAuth callback is that provider's own fixed backend URL, which this deployment has to explicitly trust. MCP_GATEWAY_REDIRECT_ALLOWLIST is that trust list: comma-separated, exact-string match (never a host or scheme — that would just move the attacker-controlled-redirect_uri hole this whole check exists to close from "any host" to "any path on one host"). It's a normal, editable setting in the anansi-app Settings UI (flag_registry.py's MCP_GATEWAY_REDIRECT_ALLOWLIST), defaulting to Claude's own documented callback, https://claude.ai/api/mcp/auth_callback (confirmed from claude.com/docs/connectors/building/authentication#callback-urls). Add a provider's URL only once confirmed from that provider's own docs, or from the exact redirect_uri a real rejected connection attempt shows in this deployment's logs — never guessed.
Setup notes. MCP_GATEWAY_BASE_URL must be the full public origin including the /mcp-gateway prefix — it's embedded in the discovery documents, is what Google validates the redirect against, and is what the mount path is derived from. The gateway reuses the admin app's Google OAuth client via AUTH_CLIENT_ID/AUTH_CLIENT_SECRET at app level, so its callback URL needs adding as a second Authorized redirect URI on that client. db/migrations/0032_oauth_code_single_use.sql must be applied to the Chat DB, or token exchange fails at the last step. None of this needs a .do/*.yaml change on its own — the gateway lives inside chat-orchestrator's existing image and ships on the next redeploy of that service, same as any other code change there.
Every setup below points at the same URL, which is specific to your deployment, not a fixed value this repo ships:
https://your-app.example.com/mcp-gateway/mcp
(substitute your own MCP_GATEWAY_BASE_URL). Every client authenticates with the connecting user's own Google account — nothing to paste, no shared secret, no per-client server-side setup beyond what's already covered above. What differs per client is purely how that client discovers a remote MCP server, split along the same CLI-vs-hosted line the OAuth section above already draws.
Claude Code (CLI — loopback, no configuration needed on this deployment):
claude mcp add --transport http anansi-mcp https://your-app.example.com/mcp-gateway/mcpApprove the browser sign-in prompt once; the token is then valid until it expires (30 days) or is revoked.
Claude Desktop / claude.ai (hosted — native remote-MCP support): Settings → Connectors → Add custom connector, paste the same URL, sign in when prompted. This is the one client path this deployment has to explicitly trust in advance: a hosted client's OAuth callback is that provider's fixed backend URL, not something the user or this server controls, so it only works if that URL is in MCP_GATEWAY_REDIRECT_ALLOWLIST (see above) — which ships with Claude's own documented callback pre-populated, so this works out of the box on a fresh deployment.
ChatGPT Desktop (bundled Codex): as of this writing, the desktop app's own MCP surface (Plugins → MCPs → Add) and Codex CLI's mcp subcommand read and write the same local config — and neither yet has native support for a remote HTTP+OAuth server the way Claude Code/Desktop do (codex mcp add only launches a local command; confirmed from its own --help, not assumed). The working path today is the community stdio↔HTTP bridge, mcp-remote (real, MIT-licensed, no telemetry — read its source before trusting a bridge with production credentials, don't take that on faith either):
codex mcp add anansi-mcp -- npx -y mcp-remote https://your-app.example.com/mcp-gateway/mcpor the same thing through the Plugins → MCPs → Add GUI: command npx, arguments -y, mcp-remote, https://your-app.example.com/mcp-gateway/mcp. mcp-remote runs its own RFC 8252 loopback OAuth flow (confirmed from its CLI usage — it takes an optional local callback port), so — like Claude Code — this needs no allowlist entry: the bridge is the thing this deployment's gateway actually sees, and it presents a loopback redirect like any other native client. If OpenAI ships true native remote-MCP support later (or if ChatGPT's own web Settings → Connectors, untested here, already does — that's a materially different, hosted-OAuth path from the desktop app's Plugins/MCPs feature, not yet verified against this gateway), that client's real callback URL would need adding to MCP_GATEWAY_REDIRECT_ALLOWLIST first, the same way Claude's was — never guessed; add only a URL confirmed from that provider's own docs, or from the exact redirect_uri a real rejected connection shows in this deployment's logs.
A client that connects successfully gets the same staff.system/customer.system prompt and knowledge modules the Telegram bot uses (see Domain context above) via InitializeResult.instructions, plus a get_operating_context tool carrying identical content for any client whose model doesn't act on that field — confirmed necessary, not speculative: an A/B test against this exact gateway showed one client resolving a site name immediately from instructions and another, receiving the identical payload, with no sign of having used it at all.
By default, App Platform builds each service's Dockerfile from scratch on
every deploy (no cross-deploy layer cache, and both services rebuild
even if only one changed) — this is what makes the default path take
several minutes. .github/workflows/build-images.yml builds and pushes
both services to GHCR with GitHub Actions' own build cache, only rebuilding
services whose paths actually changed. .do/app.image.example.yaml shows
the equivalent app spec using image: sources instead of github: +
dockerfile_path:, so App Platform just pulls a ready-made image instead
of building one.
This is opt-in and additive — the default github:-based spec above is
unaffected, and switching an existing app over is a deliberate, manual step:
# 1. Back up your current live spec first (existing safe pattern above)
doctl apps spec get <app-id> > .do/spec-backup-$(date +%Y%m%d).yaml
# 2. Base your new spec on the backup, swapping just the service `github:`/
# `dockerfile_path:` blocks for the `image:` blocks in
# .do/app.image.example.yaml (env vars, ingress, health checks, domains
# all stay the same — only the build source changes)
doctl apps update <app-id> --spec /tmp/new-spec.yaml
# 3. GHCR has no push-to-deploy webhook to App Platform (unlike DOCR), so a
# new image push doesn't auto-redeploy — trigger it explicitly:
doctl apps create-deployment <app-id>Rolling back if an image-based deploy fails: App Platform keeps your last
10 successful deployments and can restore app spec + code in one step —
click Rollback next to a prior deployment in the app's Activity tab
(or the POST /v2/apps/{app_id}/rollback API). To fully revert to the
default build-from-source path, re-apply your spec backup from step 1:
doctl apps update <app-id> --spec .do/spec-backup-<date>.yaml.
Images published to GHCR by this workflow are plain OCI images — they're
also deployable to any other container platform (Kubernetes, ECS, Fly.io,
Render, plain docker run), not just DigitalOcean.
cd chat_orchestrator
docker build -t anansi-orchestrator .
docker run -p 8000:8000 --env-file .env anansi-orchestratorSome features require external data files or third-party services. All are optional — the core chat orchestrator works without them.
The layout engine and community detection expert use GRID3 settlement-extents datasets to detect community boundaries around GPS points. GRID3 publishes these per-country for all of sub-Saharan Africa, so coverage is driven by a manifest rather than a single hardcoded file: an anchor's country is reverse-geocoded and matched against the datasets you have on hand.
Set up the data location:
- Go to https://grid3.org/resources/datasets and download the "Settlement Extents" GeoPackage (
.gpkg) for each country you operate in (e.g.NGA, ≈3.4 GB). - Put them in one directory (local path or
s3://prefix) and pointSETTLEMENT_DATA_DIRat it. - Add a
manifest.jsonin that directory describing each dataset:
{
"datasets": [
{
"iso2": "NG",
"iso3": "NGA",
"country_name": "Nigeria",
"file": "GRID3_NGA_settlement_extents_v04_3.gpkg",
"layer": "main_GRID3_NGA_settlement_extents_v4_0",
"building_count_col": "building_count"
}
]
}SETTLEMENT_DATA_DIR=/path/to/settlement-data # holds the .gpkg files + manifest.jsoniso2is the ISO 3166-1 alpha-2 code returned by reverse-geocoding (this is the match key).filemay be a bare filename (resolved againstSETTLEMENT_DATA_DIR), an absolute path, or ans3://URI — so a small local manifest can point at large remote GeoPackages.layer/building_count_colvary by country/version; copy them from each GeoPackage.- For container/S3 deploys you can instead set
SETTLEMENT_MANIFEST_JSONto the inline manifest JSON.
Adding a country: drop its .gpkg in the location and add a manifest entry — no code change.
Legacy mode: if SETTLEMENT_DATA_DIR is unset but GRID3_GPKG_PATH points at a single Nigeria GeoPackage, that still works (Nigeria-only). If neither is configured, community detection steps fail with a clear, human-readable error and the rest of the system is unaffected. Anchors in a country with no dataset get a message naming the country and listing supported ones.
The customer server includes tools for meter commissioning, unassignment, power limit control, and token resend. These call the Metering Platform API — an external meter management backend (see METERING_* env vars).
Without these env vars, meter write tools return a "not configured" error and all read-only tools continue to work normally:
METERING_API_URL=https://your-metering-instance
METERING_BEARER_TOKEN=your-bearer-token
METERING_API_KEY=your-api-keyGrid status and inverter data tools use the Victron VRM API:
VRM_API_KEY=your-vrm-api-key- Open Telegram and message @BotFather.
- Send
/newbotand follow the prompts to choose a name and username. - Copy the API token BotFather gives you — this is your
TELEGRAM_BOT_TOKEN. - Set
TELEGRAM_BOT_USERNAMEto your bot's @handle (without the@).
After deploying (or when testing locally with a tunnel like ngrok):
curl -X POST "https://api.telegram.org/bot<TELEGRAM_BOT_TOKEN>/setWebhook" \
-d "url=https://yourapp.example.com/chat" \
-d "secret_token=<API_KEY>"API_KEY must match the API_KEY env var. To verify:
curl "https://api.telegram.org/bot<TELEGRAM_BOT_TOKEN>/getWebhookInfo"Re-run setWebhook after redeployments if the URL changes.
Anansi is the single author of ticket updates in Telegram — for Jira tickets and internal (Jira-less) ones alike. Jira notifies Anansi of what changed; Anansi decides what, where, and whether to post. Turn off any direct Jira → Telegram integration (a native Jira app, an Automation rule, or an n8n flow that posts to Telegram on its own) before enabling this, or every ticket change gets announced twice.
1. Set the shared secret alongside the other JIRA_* variables above:
JIRA_WEBHOOK_SECRET=<a long random string>The endpoint is fail-closed: with no secret configured it rejects every request rather than accepting unauthenticated ones.
2. Create the webhook in Jira: Settings → System → Webhooks → Create.
| Field | Value |
|---|---|
| URL | https://yourapp.example.com/webhook/jira |
| Secret | the same JIRA_WEBHOOK_SECRET value |
| Issue events | Issue updated, Comment created |
| JQL filter | project = OPS (match your JIRA_PROJECT_KEY) |
Both event types are required — this is not a "pick one" setting:
- Issue updated — status transitions. Closure is detected via Jira's
statusCategory(done), not status names, so a custom workflow status like "Resolved" or "Completed" works without any extra configuration. - Comment created — public ("Reply to customer") comments are relayed to the escalation group, forwarded to the customer when exactly one organization matches, and mirrored into the canonical ticket-comment log so closing summaries can read Jira and internal ticket history from one place. The bot's own comments are filtered out by author email, so no reply loop forms.
Jira Cloud signs the request body with HMAC-SHA256 and sends the digest as
X-Hub-Signature: sha256=<hex>; Anansi verifies it with a constant-time
comparison and returns 401 on a mismatch.
What gets posted. On a status transition, and on any comment an LLM judges operationally significant (a diagnosis, a root cause, a blocker, a resolution — not routine chatter), Anansi renders a ticket update card — reference, status, summary, and a short summary of recent activity — and places it against that ticket's own Telegram message: edited in place while the message is still on screen, or posted as a fresh reply once the chat has moved on. Internal tickets get the identical card from the same code path, so nothing about this changes if you later turn Jira off entirely.
Verifying it works. After saving the webhook, transition a test issue and
tail the orchestrator logs for ticket update and Jira webhook lines. An
HMAC mismatch warning means the secret differs between Jira and your env
vars. Silence on a real transition usually means the issue has no active
escalation mapping — expected for a ticket Anansi never filed.
Set all required env vars in your platform:
- DigitalOcean: App Settings → Environment Variables
- Docker:
--env-fileor-eflags - Kubernetes: ConfigMaps + Secrets
Critical: Never commit .env files with real credentials!
Install pre-commit hooks for code quality checks:
# Install pre-commit
pip install pre-commit
# Install hooks
pre-commit install
# Run manually on all files
pre-commit run --all-filesEnabled checks (from .pre-commit-config.yaml):
- ✅ ruff check - Linting (rules configured in the root
pyproject.toml) - ✅ test-wiring - Fails the commit if a test file under any
tests/directory isn't tracked and wired into a CI job (see CONTRIBUTING.md "Adding a new test file")
Configured but not currently enforced by pre-commit or CI — available to run manually, settings live in the root pyproject.toml:
ruff format- Code formatting (100 line length, Black-compatible)mypy- Type checking (ignore_missing_imports = true)
detect-secrets is not wired in either; there is no .secrets.baseline in the repo.
Configuration:
.pre-commit-config.yaml- Hook definitions and pinned versionspyproject.toml- ruff and mypy settings
Common utilities in shared/:
from shared.utils.google_auth import get_drive_credentials
from shared.utils.gdrive_doc_fetcher import fetch_google_doc_markdown_sections
from shared.utils.logging import get_loggerRun ./setup_shared.sh to make shared code importable.
anansi/
├── chat_orchestrator/ # Main orchestrator
│ ├── orchestrator/
│ │ ├── api/ # FastAPI endpoints
│ │ ├── services/ # Core logic
│ │ │ ├── conversation.py # Conversation orchestration
│ │ │ ├── instructions_provider.py # Composes prompts into instructions + context
│ │ │ └── artifacts_provider.py # Section parsing
│ │ └── clients/ # External API clients
│ └── local_server.py # Development server
│
├── anansi_app/ # NiceGUI admin UI
│ ├── nicegui_app/ # Pages, layout, auth, branding (chat, grid design, Skills, settings, tickets, ...)
│ ├── grid_app/ # Grid design entities and permission helpers
│ ├── services/ # Business logic (broadcast, scheduling, Skills builder)
│ ├── scripts/ # Background jobs (broadcast_scheduler.py, grafana_scheduler.py)
│ └── db/ # Admin-app-local schema and migrations
│
├── mini_app/ # Vite customer chat widget (embedded in operator portals)
│
├── rag_pipeline/ # Knowledge ingestion
│ ├── ingestion/ # Source-specific indexers
│ └── database/ # SQL schemas
│
├── mcp_servers/ # Tool servers
│ ├── servers/ # Individual MCP servers
│ │ ├── jira_server/ # JIRA integration
│ │ ├── meters_server/ # Meter operations
│ │ ├── customer_server/# Customer/grid info
│ │ ├── grid_design_server/ # Grid design & BOM generation
│ │ ├── schedule_server/# Command scheduling
│ │ └── meta_server/ # Bot analytics
│ │ # + equipment_control, equipment_diagnostics, grafana,
│ │ # payment_processor, reference, solar, knowledge servers
│ └── mcp_launcher.py # Server manager
│
└── shared/ # Common utilities
├── prompts/
│ ├── library/ # Bundled .prompt files (the versioned default)
│ ├── core.py # PromptLibrary: resolution, render, propose/publish
│ ├── overrides.py # DB-backed versions + labels
│ ├── access.py # Per-prompt group ACLs (view/edit/publish)
│ └── knowledge.py # Tagged knowledge modules, pinned + on-demand
└── utils/
├── google_auth.py # Google authentication
├── gdrive_doc_fetcher.py # Doc fetching + parsing
└── logging.py # Logging setup
- Check the source badge on the Prompts page (
/prompts) — Default, Overridden, or Google Doc — to see where it's actually resolving from - A DB override (Overridden) takes effect within about a minute (label cache TTL); use "Reload cache" on the prompt's detail dialog to force it immediately
- A Google Doc attachment is cached for up to an hour; same "Reload cache" action applies
- Verify doc is shared with service account email
- Check service account has Docs API enabled
- Confirm doc ID is correct
- The prompt still works either way — it falls back to the bundled default or a DB override
- Add "System Instructions" as Heading 1 (not bold text)
- Section name is case-insensitive
- Must be the first section (after title page)
- Applies whether the section lives in a Google Doc or the prompt body edited from the Prompts page
./setup_shared.sh
export PYTHONPATH=$PWD# Test credentials
python3 -c "from shared.utils.google_auth import verify_credentials; verify_credentials()"- ✅ Every prompt in one place (
shared/prompts/library/), bundled and versioned with the code - ✅ Live editing from the Prompts admin page — no redeploy, no Google Doc required
- ✅ Draft → publish workflow with version history and one-click revert to default
- ✅ Per-prompt access control (view/edit/publish) via
PROMPT_EDITORS_OPS/PROMPT_EDITORS_ENG/PROMPT_ADMINS - ✅ Google Doc attachment still supported per prompt, for teams that prefer editing there
- ✅ Provenance (prompt id, source, version) on every render, in logs and Langfuse traces
- ✅ Context modules grouped by source — Built-in (code-generated), Curated (typed in the admin UI), External (attached Google Doc/Sheet) — each scoped (everywhere or one organization) and pinned explicitly per prompt or skill
- ✅ A pinned module is inlined into that prompt in full on every turn — no separate on-demand fetch step;
get_knowledge_moduleremains available as a by-name lookup tool - ✅ An attached document's audience is explicit (mirror its own Drive sharing, or publish to everyone the prompt serves) and checked against the viewing operator's own Drive access before it can even be saved
- ✅ System instructions in
systemInstructionfield - ✅ Context messages as first user message
- ✅ Token-efficient structure
- ✅ Persistent instructions across turns
- ✅ Customer mode for external users (public-facing, no sensitive data)
- ✅ Staff mode for internal users (restricted access, full knowledge)
- ✅ Identical processing pipeline
- ✅ Organization-based routing
- ✅ Multi-source ingestion
- ✅ Entity and relationship extraction
- ✅ Semantic search
- ✅ Hybrid (vector + full-text) retrieval
- ✅ Agentic graph query tools
- ✅ Community detection
- ✅ Incremental sync
- ✅ Schedule commands like
/ticketsor/gridfor future execution - ✅ Recurring schedules (daily at 9am, every monday at 10am, etc.)
- ✅ Natural language time parsing with timezone configurable via
DEFAULT_TIMEZONEenv var - ✅ Results posted to originating chat
- ✅ Staff only access (not available to customers)
- ✅ Performance reports with response vs escalation breakdown
- ✅ Escalation reason analysis with pie charts
- ✅ Negative feedback tracking
- ✅ Organization filtering for multi-tenant analytics
- ✅ Staff only access via
/metacommand
- ✅ Inline keyboard buttons supplement text-based options
- ✅ Expert workflow buttons for duplicate detection and resume prompts
- ✅ Procedure buttons for customer support conversation flows
- ✅ User mentions with Telegram deep links in group chats
- ✅ Authorization: original user, staff, or anyone for cancel
- ✅ Feature flags:
INLINE_BUTTONS_ENABLED,PROCEDURE_BUTTONS_ENABLED - ✅ Text input always works as fallback
curl -X POST http://localhost:8000/v1/chat \
-H "Content-Type: application/json" \
-d '{
"user_input": "What are your business hours?",
"user_context": {
"user_email": "user@example.com",
"source": "web"
}
}'Response:
{
"final_text": "Our business hours are Monday-Friday, 9 AM - 5 PM EST.",
"tool_calls": [],
"tool_results": [],
"history": [...]
}curl http://localhost:8000/health# Check instruction loading
grep "Loaded.*instructions" chat_orchestrator/logs/app.log
# Check RAG queries
grep "RAG retrieval" chat_orchestrator/logs/app.log
# Check errors
grep "ERROR" chat_orchestrator/logs/app.logSee CONTRIBUTING.md for the full workflow. Quick links:
- Adding an MCP Server — create a tool server end-to-end
- Expert Workflows — build multi-step LLM workflows
General guidelines:
- Follow existing code structure
- Use shared utilities from
shared/ - For prompt wording changes, edit from the Prompts admin page (no code needed); for a new prompt, add a
.promptfile undershared/prompts/library/(see CONTRIBUTING.md) - Add tests for new features
- Keep documentation current
Mozilla Public License 2.0 — see LICENSE for full text.
Questions? Check the troubleshooting section above or review the logs for detailed error messages.











