Many areas of the Philippines have slow, intermittent, or no internet access, but SMS coverage is near-universal. Pare gives those users access to an LLM assistant over plain SMS: text a question, get an answer back as a text.
Pare should feel like Google for people who are offline — one number you can text any question to, and get a useful answer back. That means general questions (math, definitions, how-to, established facts) are in scope, not just whatever happens to be in the curated corpus, and that time-sensitive answers should be genuinely up to date rather than whatever the model memorized during training.
This is deliberately more ambitious than a closed-book FAQ bot, and the two halves have very different costs. General-knowledge answers are free — they come from the model and only needed a prompt change. Up-to-date answers are not solved yet: nothing in the current architecture refreshes time-sensitive data, so the prompt forbids the model from answering those from memory (see Freshness, below).
- Reachable from any basic phone over SMS — no app, no data connection required.
- Answer broadly: retrieval-grounded where the vector store covers the topic, model general knowledge where it doesn't.
- Never present stale information as current — a refusal beats a confidently wrong answer about news, prices, or schedules.
- Every reply fits in a single SMS segment to keep per-message cost predictable (Semaphore auto-splits messages over 160 ASCII chars into multiple billed segments).
- Run at near-zero fixed cost: serverless end-to-end, pay-per-request LLM inference, no idle infrastructure.
- Multi-turn conversation memory / session state across messages.
- Rich media (MMS, images) — SMS text only.
- Multi-language UI beyond the model replying in whatever language/mix (English/Tagalog/Taglish) the user texted in.
- Authentication/user accounts — the phone number is the only identity.
- People in low-connectivity areas of the Philippines who have a phone capable of SMS but unreliable or no mobile data/wifi.
- Access is purely by sending an SMS to a published number; no signup flow.
- Current scale: personal project, roughly 10 people using it at once. Nothing here is throughput-constrained — Lambda, DynamoDB, API Gateway, and Bedrock are all pay-per-request and effectively free at this volume. The design pressure is fixed cost and maintenance burden, not scaling.
Resolved. The system previously ran retrieval through Bedrock Knowledge Bases over OpenSearch Serverless, which bills OCU-hours around the clock whether anyone texts or not — on the order of hundreds of USD/month at the standard minimum, wildly out of proportion to ~10 users' traffic. That was the whole bill for a personal project.
Retrieval was rebuilt as DIY RAG over DynamoDB: a pare-vectors table (PAY_PER_REQUEST) holds embedded chunks; src/handler.py embeds the incoming question (Bedrock Cohere Embed Multilingual) and brute-force scores every stored vector by cosine similarity in Lambda memory, no ANN index involved. At a few hundred chunks this is fast and effectively free — DynamoDB on-demand billing has no idle floor, unlike OpenSearch Serverless. Every component in the system is now pay-per-request with zero always-on cost.
The trade-off is that content ingestion is no longer fully managed: Bedrock Knowledge Bases used to handle chunking/embedding automatically on an S3 upload + sync. That's now DIY too — src/ingest.py embeds and upserts live content on its existing schedule, and scripts/ingest_corpus.py (run locally, not deployed) does the same for the curated corpus whenever it changes. More code to maintain, but no idle bill.
Considered and rejected: OpenSearch Serverless NextGen collections (GA 2026-05-28) offer true scale-to-zero and would have fixed the same problem while keeping Bedrock Knowledge Bases. As of a report dated 2026-06-02, though, Bedrock KB's Retrieve/RetrieveAndGenerate API 403s against NextGen collections (a SigV4 checksum mismatch bug), and the console's KB wizard still provisions classic collections, not NextGen — reaching NextGen requires hand-wiring a KB to an existing collection, which is exactly the path that hits the bug. If that's since been fixed, it's worth another look, but the DynamoDB approach here doesn't depend on Bedrock Knowledge Bases at all, so it isn't blocked on AWS shipping a fix.
Nothing is deployed yet (see Current state), so this is a projection, not a bill — but every component is pay-per-request with no idle floor, so the total is small at this project's scale regardless of exact traffic. Assumes the default hourly IngestSchedule (template.yaml) with a moderate source mix (a handful of weather cities + PAGASA + one news feed) and a few hundred SMS/month, well under the DAILY_LIMIT=15-per-number ceiling.
| Component | Driver | Est. monthly cost |
|---|---|---|
| Lambda (webhook + ingest) | ~1,200 invocations/mo, 256 MB | $0 — inside the always-free tier (1M req + 400K GB-s) |
| API Gateway (HTTP API) | ~500 webhook calls/mo @ $1/M requests | ~$0.001 |
| EventBridge Scheduler | 720 hourly ingest triggers/mo | $0 — inside the 14M free invocations/mo |
DynamoDB (pare-vectors + pare-rate-limits, PAY_PER_REQUEST) |
scan-per-message reads + hourly upsert writes over a few hundred chunks | ~$0.10–0.50 |
| Bedrock — Nova Lite generation | ~500 msgs × (~500 in / ~50 out tokens) @ $0.06 / $0.24 per 1M tokens | ~$0.02 |
| Bedrock — Cohere Embed Multilingual | query embeds + hourly ingest re-embeds of ~15 chunks | ~$0.15–0.30 |
CloudWatch Alarms (3, template.yaml) |
fixed per-alarm charge | ~$0.30–0.50 |
| CloudWatch Logs | JSON logs, low volume | ~$0.01–0.05 |
| SSM Parameter Store | 1 Standard-tier SecureString | $0 |
| Total | ~$1–3/month, comfortably under $5 even with margin |
For scale, the OpenSearch Serverless floor this design replaced ran on the order of $350+/month minimum, regardless of traffic — the whole motivation for the rebuild described above. Even pushing ingest to cover all 32 WEATHER_CITIES entries hourly, or maxing out DAILY_LIMIT across 10 numbers (4,500 msgs/month), keeps the total under ~$5/month; Bedrock's per-token rates are cheap enough that traffic within this project's stated scale can't meaningfully move the bill. The one-time cost of running scripts/ingest_corpus.py locally (your own AWS credentials, same Bedrock/DynamoDB rates) is trivial for a few dozen documents and isn't included above.
- SMS gateway is Semaphore, a Philippine SMS provider — not Twilio. Its v4 REST API, auth, and payload shapes are different from Twilio conventions (no TwiML, no
twilioSDK). See system-design.md for verified details. - Reply budget: 150 characters, enforced in both the system prompt (soft) and code (hard truncation), to always stay inside Semaphore's single-segment (160 ASCII char) threshold with headroom.
- Semaphore silently drops any message starting with "TEST" — handled defensively in code (see
_fit_for_smsinsrc/handler.py). - DIY RAG over DynamoDB: the project has gone through two RAG architectures now. It started with manual RAG (external vector DB + NVIDIA NIM via the OpenAI SDK), moved to Amazon Bedrock Knowledge Bases (S3 + OpenSearch Serverless, one
RetrieveAndGeneratecall), and moved again to the current design once OpenSearch Serverless's always-on OCU floor turned out to dominate the bill at this project's scale (see Cost posture). Retrieval and generation are now two separate Bedrock calls: embed the question (Cohere Embed Multilingual viainvoke_model), brute-force cosine similarity over every vector in thepare-vectorsDynamoDB table, thenconverse(Nova Lite) with the top matches spliced into the prompt. This trades Bedrock Knowledge Bases' managed chunking/embedding/indexing for DynamoDB's pay-per-request pricing with no idle floor — the same embedding-model coupling between ingest and query that a fully-managed KB avoided is now somethingsrc/handler.py,src/ingest.py, andscripts/ingest_corpus.pyall have to agree on by convention (sameEMBED_MODEL_ID, same asymmetricinput_typecontract). - Generation model is Amazon Nova Lite (on Bedrock), called via the plain
converseAPI. The project originally specified Qwen3 32B, then switched to Nova Lite because Qwen3 was excluded from Bedrock Knowledge Bases' supported-model list. That specific constraint no longer applies now that generation doesn't go through Knowledge Bases, but Nova Lite remains the committed choice — it's cheap and needs no cross-Region inference profile. Whatever the model, the Lambda has zero third-party dependencies — noopenaioranthropicSDK; everything AWS goes through boto3. - Tagalog quality is an open risk — AWS doesn't list Tagalog among Nova Lite's explicitly-optimized languages. Nova is the committed choice regardless, so if replies read poorly the lever is the prompt template and knowledge-base content, not a model swap.
- Embedding model is Cohere Embed Multilingual, forced by regional availability: Titan Text Embeddings V2 isn't offered in
ap-southeast-1at all. Cohere covers 100+ languages, so this constraint happens to suit Tagalog content. Called directly viabedrock-runtime.invoke_modelnow (no longer bound to a Knowledge Base at creation time) —EMBED_MODEL_IDis a plain constant duplicated acrosssrc/handler.py,src/ingest.py, andscripts/ingest_corpus.py, and all three must agree on both the model and theinput_typeconvention (search_queryvssearch_document) or retrieval quality silently degrades. - Inbound webhook payload field names from Semaphore are not publicly documented, so the parser accepts several plausible key names for the sender number and message text and needs to be pinned down once a real webhook payload is observed in production logs.
The model's knowledge is frozen at its training cutoff, and the vector store only contains whatever was last embedded and written. The prompt template contains the risk rather than solving it: time-sensitive categories (news, prices, weather, schedules, current officials, sports) may only be answered from search results, and Pare says it has no updated info otherwise. Honest, but it means a refusal unless something is actively keeping those search results fresh.
Scheduled ingest (src/ingest.py) is that something. An EventBridge schedule invokes a second Lambda that fetches configured RSS/Atom feeds, embeds each rendered body (Bedrock Cohere Embed Multilingual), and upserts it into the pare-vectors DynamoDB table. The webhook Lambda is unaware of it — it scans and scores whatever is currently in the table. Freshness is bounded by the schedule interval, not by the model.
Two invariants make this safe, and both are enforced by tests:
- Fixed item keys, overwritten in place. Writing dated keys instead would grow the table forever and let a week-old headline out-rank today's — cosine similarity has no notion of recency on its own.
- Each item's text carries its own "current as of" stamp and fits in one embedding call, so a retrieved advisory can never be separated from its timestamp.
Three source types are wired up, chosen so that neither weather source needs an API key or a second secret:
- Open-Meteo (JSON, no key, 1–11 km resolution) — one file per city from a
CITIEStable: all 17 NCR LGUs plus 15 provincial centers across Luzon, Visayas, and Mindanao. Hyperlocal without any location state in the handler: each file repeats its city name, so "ulan ba sa Cebu?" retrieves the Cebu chunk by embedding match alone. NCR is covered LGU-by-LGU so a texter's own city is named back at them, accepting that adjacent Metro Manila LGUs will often report identical conditions — the region is barely wider than the model's grid. - PAGASA advisory page (server-rendered HTML, scraped) — tropical cyclone wind signals. No global weather API carries PAGASA's Signal No. 1–5, and that is the fact that matters most during a storm. PAGASA's ten-day forecast API exists but is token-gated; the advisory page is not.
- RSS/Atom news — URL-configured, off by default. Inquirer, GMA, PhilStar, and Rappler feeds are verified working (see
deployment.md§17a for URLs);abs-cbn.comandpna.gov.phboth return 403 to automated clients and cannot be used. The Philippine News Agency would have been the preferred source — government wire, cleanest redistribution story — but it blocks automated fetches, so the working options are all commercial publishers.
What this still doesn't cover: anything outside the feeds you configure. A question about a topic with no feed behind it gets the refusal path, which is the correct outcome. Closing that gap generally would mean live web search at query time — replacing the DynamoDB retrieval step with a generation call against a provider offering search grounding (e.g. the Gemini API's Grounding with Google Search). That's the freshest option but moves generation off Bedrock, adds a non-AWS dependency and a second API key, and breaks the console-only deployment story in deployment.md.
Because the phone number is the only identity and every message costs real money (Bedrock embedding + generation tokens, outbound SMS credits), each number is capped at a daily message limit (default 15, via DynamoDB atomic counters with TTL). Over-limit senders get a polite static reply and no Bedrock spend occurs. The limiter fails open: if DynamoDB errors, the bot keeps answering.
The core Lambda handler, ingest Lambda, corpus ingestion script, SAM infrastructure template (with DynamoDB rate-limit table, DynamoDB vector table, and CloudWatch alarms), and test suite exist (see src/, scripts/, template.yaml, tests/). The Semaphore API key is designed to live in an SSM Parameter Store SecureString. Not yet done: no pare-vectors table created, no curated documents run through scripts/ingest_corpus.py, no deployed stack, no SSM parameter created, and no confirmed real-world Semaphore webhook payload. See system-design.md for the architecture and open questions.