Skip to content

Repository files navigation

Screenshot 2026-08-22 at 3 37 08 PM

Paytm AI Soundbox - TEAM ERROR :) - checkout the video here : https://www.youtube.com/watch?v=EmD8BHPeV6Q

Tech Stack : SARVAM AI API - NEXT.js, Supabase, Postgresql.

The Soundbox already tells a merchant that money arrived. This one tells them what the money means. **In both telugu and English languages for rural and urban merchant customers too **

paytm-telugu.mp4
paytm-english.mp4

An AI CFO for small merchants, built into the device they already trust — ask a question out loud in Telugu or English, get a real answer computed from your own payment data.

Built for HackCulture, 22 Aug 2026. Demo merchant: Cafe Niloufer, an Irani café in Hyderabad — 86,211 orders, ~218,000 line items, 12 months of trading.

Screenshot 2026-08-23 at 12 49 01 PM

The idea in one paragraph

Every small merchant in India already has a Paytm Soundbox on the counter. It speaks — but it only ever says one thing: "₹150 received." Meanwhile the merchant has no idea that their regulars are drifting away, that evenings are quietly dying, or that revenue is up while profit is not.

Paytm AI Soundbox turns that one-way announcer into something closer to Alexa with a finance brain. The merchant asks "వ్యాపారం ఎలా ఉంది?" and the box answers — not with a dashboard, but in a sentence: what is happening, why, and what to do this week.

The one rule this project is built on

Code calculates the facts. AI interprets them.

Every number the CFO says was computed by SQL over 217,000 order lines. The model never sees a raw row and never does arithmetic. It receives small, already- correct aggregates and reasons about what they mean.

That is the entire difference between an AI CFO and a chatbot pointed at a dashboard — and it is why the answers can be trusted on stage.


What it found (the demo story)

The dataset looks healthy on the surface. Revenue is up 29% year on year.

The CFO says otherwise:

What the owner sees What is actually happening
Revenue +29% YoY Growth is bought with one-time footfall
"Returning customers +57%" That metric is lying — acquisition inflates it
— Orders per active customer 4.50 → 3.96
— Loyal customers' share of revenue 81% → 74.5%
— Evening trade (5–8pm) 29.2% → 27.2%
— Average order value ₹165 → ₹150

The regulars are leaving, and new footfall is hiding it.

That naive "returning customer count" is the most important thing in this project. It rises while the business rots, because every new customer who comes twice counts as "returning". The tool descriptions warn the model about it explicitly so it can never pick the flattering metric over the honest one.


The three surfaces

Route What it is
/ Paytm-style storefront — the shell the product lives in
/dashboard Merchant console: overview, analytics, menu, customers, transactions
/cfo Chat as a particle field — cursor wakes the data, chips form from what it touches
/voice The Soundbox, talking — speech in, speech out, ringed by a live audio visualiser

/voice is the point of the whole thing. The Soundbox sits at the centre of a 128-bar radial visualiser and breathes while it speaks, so the device reads as the thing answering you.


How it works

     microphone
         │  Web Speech API  (free, unlimited, te-IN)
         ▼
    ┌──────────────────────────────────────────────┐
    │  SARVAM  sarvam-105b-conversations           │
    │  picks tools · interprets · never calculates │
    └──────────────────────────────────────────────┘
         │  12 tools + get_full_picture
         ▼
    ┌──────────────────────────────────────────────┐
    │  POSTGRES  (Supabase)                        │
    │  12 analytics functions over a materialised  │
    │  view of 217k order lines                    │
    └──────────────────────────────────────────────┘
         │  small, exact aggregates
         ▼
    Sarvam Bulbul v3 TTS  (native te-IN)
         │
         ▼
     the Soundbox speaks

Stack

Layer Choice
Data Python + Faker + Pandas — seeded generator, 8 behaviour rules
Database Supabase Postgres, materialised view + 12 SQL functions
Backend Node + TypeScript + Express, 16 REST endpoints
AI Sarvam (sarvam-105b-conversations), Gemini + Groq as failover
Voice Sarvam Bulbul v3 TTS, ElevenLabs as fallback; browser STT
Frontend Next.js + React + Tailwind + Recharts

The data is generated, and that is the point

No CSV was hand-edited. dg/config/business_config.py holds 8 business behaviour rules — each a probability multiplier that ramps over time — and the simulator runs the café day by day for a year. The patterns emerge from thousands of individual customer decisions.

Rule What changes What the AI finds
evening_slowdown 5–8pm order probability × 0.70 evening share 31% → 27%
churn_cohort 55% of regulars decay to ~0 orders per customer −28%
falling_aov basket size × 0.90 AOV ₹186 → ₹155
margin_squeeze chai + bakery cost × 1.15 gross margin slips
declining_product Mango Lassi × 0.55 share −37%
dead_stock 4 items ≈ 0 4 items under 0.4% of volume
weekend_spike Sat/Sun × 1.30 healthy rhythm (context)
rising_star Irani Chai × 1.38 share +28%

PATTERNS_ENABLED = False regenerates a clean baseline, which is how we prove a pattern is caused by its rule and not by chance.

cd dg && pip install numpy pandas faker
python3 main.py                 # patterns ON   -> data/       (~30s)
python3 main.py --no-patterns   # baseline      -> data_baseline/
python3 -m validate.checks      # all 9 checks pass

⚠️ Primary keys use uuid4(), which is not seeded. A rerun is statistically identical but structurally different — so a regenerated dataset is all-or-nothing: truncate and reload all 7 tables, then re-run analytics.sql to rebuild the materialised view.


Latency: 41s → 10–20s

The honest breakdown, because it drove most of the engineering:

Before After
Full answer 41.0s 10–20s
Cached / prewarmed 35.1s instant
  1. Measured first. Instrumented every answer with modelMs / toolMs / rounds. The database was only 2–5s — ~90% was model thinking.
  2. Killed a latch bug. A generic 500 was silently disabling thinking control process-wide, reverting everything to HIGH thinking. 41s → 21.6s on its own.
  3. Split effort by round. Choosing a tool is pattern matching; only the final answer round gets real reasoning.
  4. get_full_picture. One tool runs all 12 analyses concurrently and returns a compacted digest — collapsing several sequential round trips into one. Faster and more thorough.
  5. Disk answer cache + prewarm. Demo questions cost zero API calls.

What did not work: lowering the max tool rounds. The loop exits the moment the model stops asking for tools, so an unused round costs nothing — it only truncated the hard questions. Reverted.


Running it

# backend  :4000
cd backend && npm install && cp .env.example .env   # fill in the keys
npm run dev

# frontend :3000
cd frontend && npm install && cp .env.local.example .env.local
npm run dev

Health checks:

npm run smoke                                  # 14/14
curl localhost:4000/api/ai/ping    | python3 -m json.tool   # which AI engine, and is it alive
curl localhost:4000/api/voice/ping | python3 -m json.tool   # which TTS engine, and does Telugu work

Two things that will otherwise cost you an afternoon:

  • .env does not hot-reload under tsx watch. Code does. Change .env → restart. /api/ai/status prints processStartedAt so you can see it.
  • Google's free-tier quota is per project, not per key. A new API key in the same project draws from the same empty bucket.

See TEAMMATE-SETUP.md for the full setup, and journal.txt for the complete build log — every decision, every bug, and why.


Design notes worth knowing

  • Merchants are not analysts. The CFO never says "cohort retention declined 19.3%" — it says "your regulars are coming in less often", then gives the number. The prompt enforces this.
  • The owner's language choice is authoritative. Speech recognition set to English will happily transcribe spoken Telugu as English-looking words, so the UI toggle overrides whatever the text appears to be.
  • Provider failover. Sarvam → Gemini → Groq for the brain, Sarvam Bulbul → ElevenLabs for the voice. A dead key degrades instead of silencing the product.
  • No names, no phone numbers. The console identifies a customer by id and an order by reference. The API returns names; the UI ignores them.
  • No network fonts, no CDN components. Everything renders from local assets — a font request that fails on stage is a demo that fails on stage.

Repository

dg/          seeded data generator + schema.sql + analytics.sql (12 SQL functions)
backend/     Express API, AI agent loop, provider abstraction, TTS
frontend/    Next.js app — storefront, console, /cfo, /voice
journal.txt  the full build log, ~1,600 lines
idea.txt     the original idea dump
blueprint.txt the 11-row plan we actually followed

Releases

Packages

Contributors

Languages