How to take Pare from this repo to a live SMS number using only the AWS Console — no AWS CLI, no CloudFormation/SAM. Architecture background: system-design.md.
template.yaml still exists in the repo as an Infrastructure-as-Code reference matching this same architecture, in case you want to compare or automate later. It is not used in this guide, and console-created resources can drift from it — don't assume the two stay in sync.
Because both Lambdas have zero third-party dependencies (boto3 ships with the runtime), their code is small enough to paste directly into the Lambda console's built-in editor — no zip files, no local build step, nothing to upload.
The stack: Amazon Nova Lite for generation (Bedrock converse API), Cohere Embed Multilingual for embeddings (Bedrock invoke_model), a DynamoDB table (pare-vectors) as the vector store — retrieval is a brute-force cosine-similarity scan done in Lambda, not a managed search service — in Asia Pacific (Singapore) ap-southeast-1.
Why DynamoDB instead of a Bedrock Knowledge Base. An earlier version of this guide used Bedrock Knowledge Bases over OpenSearch Serverless, which handles chunking/embedding/indexing automatically. OpenSearch Serverless bills an always-on OCU floor (order of hundreds of USD/month) regardless of traffic, which dominated the bill for this project's ~10-user scale. DynamoDB's on-demand billing has no idle floor — the trade-off is that you now run a small local script (
scripts/ingest_corpus.py) to embed curated documents instead of uploading them to S3 and clicking Sync. See context.md Cost posture.
Everything you'll create, in the order you'll create it. Keep this handy for teardown (§15).
Items 1–8 are the core SMS bot. Items 9–11 (§17) add the scheduled ingest that makes time-sensitive answers possible — deploy them once 1–8 work end to end.
| # | Resource | Name | Created in |
|---|---|---|---|
| 1 | DynamoDB table (vector store) | pare-vectors |
§3 |
| 2 | SSM Parameter (SecureString) | /pare/prod/semaphore-api-key |
§5 |
| 3 | DynamoDB table (rate limit) | pare-rate-limits |
§6 |
| 4 | IAM role | pare-sms-webhook-role |
§7 |
| 5 | Lambda function | pare-sms-webhook |
§8 |
| 6 | API Gateway HTTP API | pare-api |
§10 |
| 7 | CloudWatch alarms | pare-sms-webhook-error-rate, pare-sms-webhook-duration-p95 |
§12 |
| 8 | SNS topic (optional) | pare-alarms |
§13 |
| 9 | IAM role (ingest) | pare-ingest-role |
§17 |
| 10 | Lambda function (ingest) | pare-ingest |
§17 |
| 11 | EventBridge schedule | pare-ingest-hourly |
§17 |
Not a deployed resource, but part of getting real answers: §18 curated corpus ingestion, run locally with scripts/ingest_corpus.py whenever you add or change a document.
- An AWS account, signed into the Console, with permissions to create IAM roles, Lambda functions, API Gateway APIs, DynamoDB tables, SSM parameters, and CloudWatch alarms (an Administrator role is simplest if this is your own account).
- A Semaphore account (https://semaphore.co) with credits, your API key from the dashboard, and (optionally) a registered sender name.
- Your AWS Account ID: click your account name in the top-right corner of the console — the 12-digit ID is shown there with a copy icon. You'll need it several times below to build ARNs by hand.
- Local AWS credentials configured (
aws configure, or an SSO profile) for runningscripts/ingest_corpus.py— this script is not deployed, it runs on your machine.
Use the region selector in the top-right of the console and pick the same region for every step below. This guide uses Asia Pacific (Singapore) ap-southeast-1 — a Bedrock region that supports both Nova Lite and Cohere Embed Multilingual and is close to the Philippines. If you use a different region, substitute it everywhere and re-check step 2.
There is nothing to enable. Access to serverless foundation models is automatic for any account with a valid payment method. Console → Amazon Bedrock → Model catalog → search for Nova Lite and Cohere Embed Multilingual to confirm both appear for your account. Optionally hit Open in playground to sanity-check each responds — Nova Lite under Chat/Text, Cohere Embed Multilingual under Embeddings (it'll show a vector of numbers, which is expected).
If you're following an older tutorial that tells you to visit a Model access page and click Request model access, that flow no longer exists — don't go looking for it.
Console → DynamoDB → Tables → Create table:
- Table name:
pare-vectors - Partition key:
pk, type String - Table settings: Customize settings → Read/write capacity settings: On-demand
- Create table
That's it — there's no schema to define beyond the partition key. src/ingest.py and scripts/ingest_corpus.py write text, embedding, source, and (for live content) updated_at as plain item attributes; DynamoDB doesn't need to know about them in advance.
Before wiring anything else up, confirm both Bedrock calls work from the console:
- Bedrock console → Model catalog → Cohere Embed Multilingual → Open in playground. Enter a short sentence, e.g. "Ano ang kabisera ng Pilipinas?", and confirm you get back a vector of numbers (a few thousand of them). This is exactly what
_embed()insrc/handler.py/src/ingest.py/scripts/ingest_corpus.pycalls at runtime. - Bedrock console → Model catalog → Nova Lite → Open in playground. Ask a question and confirm you get a coherent reply. This is what
_generate()calls viaconverse.
If either fails here, fix it before continuing — everything downstream depends on both.
Console → Systems Manager → Parameter Store (left sidebar, under Application Management) → Create parameter:
- Name:
/pare/prod/semaphore-api-key - Tier: Standard
- Type: SecureString
- KMS key source: My current account (default
aws/ssmkey) - Value: paste your Semaphore API key
- Create parameter
Console → DynamoDB → Tables → Create table:
- Table name:
pare-rate-limits - Partition key:
pk, type String - Table settings: Customize settings → Read/write capacity settings: On-demand
- Create table
Once the table status shows Active, open it → Additional settings tab → Time to Live (TTL) → Turn on:
- TTL attribute:
expires_at - Turn on TTL
This attribute (written by the Lambda) is what lets DynamoDB auto-delete stale daily counters — no cleanup job needed.
Console → IAM → Roles → Create role:
- Trusted entity type: AWS service. Use case: Lambda. Next.
- Permissions policies: search for and check AWSLambdaBasicExecutionRole (covers CloudWatch Logs). Next.
- Role name:
pare-sms-webhook-role. Create role.
Now add the app-specific permissions. Open the role you just created → Permissions tab → Add permissions dropdown → Create inline policy → JSON tab → replace the contents with the policy below, substituting <REGION> (ap-southeast-1) and <ACCOUNT_ID> (from §0) with your real values:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "bedrock:InvokeModel",
"Resource": [
"arn:aws:bedrock:<REGION>::foundation-model/amazon.nova-lite-v1:0",
"arn:aws:bedrock:<REGION>::foundation-model/cohere.embed-multilingual-v3"
]
},
{
"Effect": "Allow",
"Action": "ssm:GetParameter",
"Resource": "arn:aws:ssm:<REGION>:<ACCOUNT_ID>:parameter/pare/prod/*"
},
{
"Effect": "Allow",
"Action": "dynamodb:UpdateItem",
"Resource": "arn:aws:dynamodb:<REGION>:<ACCOUNT_ID>:table/pare-rate-limits"
},
{
"Effect": "Allow",
"Action": "dynamodb:Scan",
"Resource": "arn:aws:dynamodb:<REGION>:<ACCOUNT_ID>:table/pare-vectors"
}
]
}Click Next, name the policy pare-sms-webhook-policy, and Create policy.
If the SSM parameter in §5 ends up encrypted with a customer-managed KMS key instead of the default
aws/ssmkey, add akms:Decryptstatement on that key too.
- Console → Lambda → Functions → Create function.
- Author from scratch. Function name:
pare-sms-webhook. Runtime: Python 3.12. Architecture: arm64 (cheaper per ms than x86_64). - Expand Change default execution role → Use an existing role → select
pare-sms-webhook-role. - Create function.
Add the code: in the Code tab, the built-in editor shows a default lambda_function.py.
- In the file explorer pane on the left of the editor, right-click → New File, name it
handler.py. - Open
src/handler.pyfrom this repo, copy its entire contents, and paste them into the newhandler.pyin the console editor. - Right-click
lambda_function.pyin the file explorer → Delete (no longer needed). - Click the orange Deploy button above the editor to save.
Set the handler: scroll down to Runtime settings → Edit → Handler: handler.lambda_handler → Save.
Set memory/timeout: Configuration tab → General configuration → Edit:
- Memory:
256MB - Timeout:
0min29sec (API Gateway's own integration ceiling is ~29s, so there's no benefit to going higher) - Save
Configuration tab → Environment variables → Edit → Add environment variable, once per row:
| Key | Value |
|---|---|
MODEL_ARN |
arn:aws:bedrock:ap-southeast-1::foundation-model/amazon.nova-lite-v1:0 |
VECTOR_TABLE |
pare-vectors |
SSM_PREFIX |
/pare/prod |
RATE_LIMIT_TABLE |
pare-rate-limits |
DAILY_LIMIT |
15 |
SEMAPHORE_SENDER_NAME |
your registered sender name (optional — omit this row entirely if you don't have one) |
Save.
Test tab → Create new test event. Event name: webhook-test. Paste this, replacing the phone number with your own real number so you can confirm delivery, not a fake number that just burns a Semaphore credit on a failed send:
{
"version": "2.0",
"routeKey": "POST /webhook/sms",
"rawPath": "/webhook/sms",
"isBase64Encoded": false,
"requestContext": { "http": { "method": "POST" } },
"body": "{\"number\":\"<YOUR_OWN_NUMBER>\",\"message\":\"Ano ang kabisera ng Pilipinas?\"}"
}Save, then click Test. You should see a successful execution result ({"statusCode": 200, "body": "OK"}) and a text message on your phone shortly after — this question is general knowledge, so it should work even with pare-vectors still empty. If it fails, expand Execution results for the error, or check Monitor tab → View CloudWatch logs.
- Console → API Gateway → APIs → Create API → find HTTP API → Build.
- Add integration → Lambda → select your region and function
pare-sms-webhook. API name:pare-api. Next. - Configure routes: Method POST, resource path
/webhook/sms, integration target the Lambda integration you just added. Next. - Configure stages: leave stage name as
$defaultwith Auto-deploy enabled. Next → Create.
This wizard automatically grants API Gateway permission to invoke your Lambda — unlike a from-scratch setup, there's no separate permission step to remember.
Copy the Invoke URL shown on the API's main page (or under Stages → $default). Your full webhook URL is that plus the route path:
https://<api-id>.execute-api.ap-southeast-1.amazonaws.com/webhook/sms
If you want to test this layer specifically before spending a Semaphore credit, a GUI HTTP client like Postman works well — POST the same JSON body from §10 to this URL and confirm you get OK back and a text arrives. This is optional; §12's real test covers the same ground.
- In the Semaphore dashboard, set the incoming-message webhook to the URL from §11. (If there's no self-serve webhook setting, contact Semaphore support to enable inbound messages for your account — inbound/webhook behavior isn't in their public docs.)
- Text a question to your Semaphore number from a real phone.
- Console → CloudWatch → Log groups →
/aws/lambda/pare-sms-webhook→ open the latest log stream. You should see structured JSON lines:inbound_smsthensms_sent. - First-message action item: if the reply doesn't arrive, look for a
webhook_missing_fieldslog entry — it includes the raw payload. Opensrc/handler.pylocally, correct_NUMBER_KEYS/_MESSAGE_KEYSto the real field names, then redeploy the code (§15, "Update code"). This is expected on first contact since Semaphore's inbound payload shape is undocumented.
For repeated log digging, CloudWatch → Logs Insights → select the log group → run:
fields @timestamp, message, number, chars
| filter message in ["inbound_sms", "sms_sent", "rate_limited"]
| sort @timestamp desc
Error rate (percentage isn't a raw metric, so this uses a math expression):
- CloudWatch → Alarms → All alarms → Create alarm → Select metric.
- Browse Lambda → By Function Name → find
pare-sms-webhook→ check both Errors and Invocations → Select metric. - Below the graph, in the metrics table, click Add math expression → Start with empty expression. Formula (use the row IDs shown for Errors/Invocations, typically
m1/m2):Label itIF(m2>0, 100*m1/m2, 0)ErrorRatePercent. Uncheck "use in alarm" (or equivalent) for the two raw metrics so only this expression drives the alarm. - Next → Threshold type Static, "whenever ErrorRatePercent is..." Greater than
1→ Next. - Notification: select or create an SNS topic (§14), or skip. Next.
- Name:
pare-sms-webhook-error-rate. Description: "Lambda error rate above 1% over 5 minutes. The handler catches its own exceptions, so this firing means crashes outside the handler: import errors, timeouts, or OOM." → Create alarm.
Duration p95:
- Create alarm → Select metric → Lambda → By Function Name →
pare-sms-webhook→ check Duration → Select metric. - Statistic dropdown: choose p95 (under Percentile). Period: 5 minutes. Next.
- Threshold: Static, "whenever Duration p95 is..." Greater than
10000(milliseconds). Additional configuration: Datapoints to alarm 2 out of 2. Missing data treatment: Treat missing data as good (not breaching). Next. - Notification: same SNS topic or skip. Next.
- Name:
pare-sms-webhook-duration-p95. Description: "p95 Lambda duration above 10s — embedding/retrieval/generation is slow." → Create alarm.
Without this, alarms fire silently into CloudWatch only.
Console → SNS → Topics → Create topic → Standard → name pare-alarms → Create topic. Then Create subscription → Protocol Email → Endpoint your email address → Create subscription, and confirm via the link AWS emails you. You can create this ahead of time or inline during the alarm wizard's notification step (§13).
- Update the curated corpus: add/edit
.txt/.mdfiles under a local directory (e.g.data/), then run.venv\Scripts\python scripts\ingest_corpus.py data— see §18. No console step needed. - Update live content: nothing manual —
pare-ingestruns on its own schedule (§17). - Update code: Lambda console → Code tab → edit
handler.py(oringest.py) in the built-in editor → Deploy. - Update config: Lambda console → Configuration → Environment variables → Edit (this fully replaces the set shown, so double-check nothing existing got dropped before saving).
- View logs: CloudWatch console → Log groups →
/aws/lambda/pare-sms-webhook, or Logs Insights for the query above. - Tear down, roughly in reverse order:
- Semaphore dashboard: remove the webhook URL first, so it stops POSTing to a dying endpoint.
- API Gateway console → select
pare-api→ Delete. - Lambda console → select
pare-sms-webhook→ Actions → Delete. - IAM console →
pare-sms-webhook-role→ detachAWSLambdaBasicExecutionRole, delete the inlinepare-sms-webhook-policy, then delete the role. - DynamoDB console →
pare-rate-limitstable → Delete. - CloudWatch console → Alarms → select both alarms → Actions → Delete.
- Systems Manager → Parameter Store →
/pare/prod/semaphore-api-key→ Delete. - SNS console →
pare-alarmstopic → Delete (if created). 8b. If you did §17: EventBridge → Schedules →pare-ingest-hourly→ Delete; Lambda →pare-ingest→ Delete; IAM →pare-ingest-role→ Delete; CloudWatch →pare-ingest-errorsalarm → Delete. - DynamoDB console →
pare-vectorstable → Delete.
Everything here is pay-per-request with no idle floor — DynamoDB on-demand billing, Lambda, API Gateway, SSM, and Bedrock are all zero-cost at zero traffic. Historically the biggest line item was Bedrock Knowledge Bases' required OpenSearch Serverless collection (order of hundreds of USD/month at the standard minimum, billed 24/7 regardless of traffic); replacing it with a DynamoDB vector table removes that floor entirely. See context.md Cost posture for the full history and the estimate this table summarizes.
At SMS volumes, Lambda/API Gateway/DynamoDB/SSM round to nothing. Bedrock and Semaphore bill per message: one embedding call plus one generation call per question, both against cheap models (Cohere Embed Multilingual, Nova Lite). The §17 ingest and §18 corpus tool barely move this either — a handful of short embedding calls a day/per-edit costs fractions of a cent, and Open-Meteo is free at this volume. Because both ingestion paths overwrite fixed keys instead of appending, the vector table doesn't grow unbounded over time.
Rough monthly total: ~$1–3, comfortably under $5 even with margin — assuming the hourly IngestSchedule from §17f, a moderate source mix (a few weather cities + PAGASA + one news feed), and a few hundred SMS/month (well under DAILY_LIMIT=15 per number). Breakdown:
| Component | Est. monthly cost |
|---|---|
| Lambda (webhook + ingest) | $0 — inside the always-free tier |
| API Gateway (HTTP API) | ~$0.001 |
| EventBridge Scheduler (§17f) | $0 — inside the 14M free invocations/mo |
DynamoDB (pare-vectors + pare-rate-limits) |
~$0.10–0.50 |
| Bedrock — Nova Lite generation | ~$0.02 |
| Bedrock — Cohere Embed Multilingual | ~$0.15–0.30 |
| CloudWatch Alarms (§13, §17g) | ~$0.30–0.50 |
| CloudWatch Logs | ~$0.01–0.05 |
| SSM Parameter Store (Standard tier) | $0 |
Pushing ingest to all 32 WEATHER_CITIES hourly, or maxing DAILY_LIMIT across 10 numbers, still keeps the total under ~$5/month — Bedrock's per-token rates are cheap enough that traffic at this project's stated scale can't move the bill much. §18's corpus tool runs on your own credentials outside this Lambda-side estimate, but bills at the same per-token/per-request rates and is trivial for a few dozen documents.
At roughly 10 users this whole system should now be near-free end to end.
Everything up to here answers from your curated corpus (§18) plus the model's general knowledge, and by design it refuses questions about news, prices, weather, and schedules (see the prompt rules in src/handler.py). This section is what makes it able to answer them — it's the half of the product that delivers up-to-date data, so treat it as deferred rather than optional.
The idea is small: a second Lambda on a timer fetches current data, embeds it, and upserts it into the same pare-vectors table the webhook scans. pare-sms-webhook doesn't change at all — it scans and scores whatever is currently in the table.
Do this only after §12 works end to end.
17a. Know your three source types. src/ingest.py ships with all three already wired:
| Source | Type | Config | Needs a key? |
|---|---|---|---|
| Open-Meteo per-city weather | JSON API | WEATHER_CITIES |
No |
| PAGASA cyclone wind signals | Scraped HTML | PAGASA_ADVISORY_URL |
No |
| News headlines | RSS/Atom | NEWS_FEED_URL |
No |
Weather and typhoon signals work out of the box — neither needs an account, an API key, or a second SSM secret. News is the only one you have to supply a URL for.
News feed URLs, tested against
ingest.pyon 2026-08-01. Fetched, parsed, and rendered end to end — all four working feeds produced output under the 1200-char embedding cap:
Feed URL Entries Notes Inquirer https://newsinfo.inquirer.net/feed40 Recommended — domestic news section, most material to render GMA https://data.gmanetwork.com/gno/rss/news/feed.xml15 Solid alternative PhilStar https://www.philstar.com/rss/headlines10 Lightest payload (7 KB) Rappler https://www.rappler.com/feed/10 131 KB payload, mixes in international tech news Blocked — do not use:
abs-cbn.comandpna.gov.phboth return 403 to automated clients. There's no working around it; pick from the table above.All four working feeds are commercial publishers, so you're storing headlines plus ~200 characters of summary from copyrighted reporting. Pare emits short derived answers rather than reproductions, which is a defensible posture at personal-project scale — but it is a choice, not a default.
Why weather is split across two sources: Open-Meteo gives hyperlocal conditions and rainfall (1–11 km resolution, lat/lon, free, no key) but publishes no PAGASA wind signals. "Signal No. 3 na ba dito?" is the question that actually matters during a storm, and only PAGASA issues that. PAGASA's ten-day forecast API is token-gated, but its advisory page is server-rendered, so
ingest.pyscrapes the headings out of it. Two sources, one weather story.
How location works without location tracking. Pare has no idea where a texter is — the phone number is the only identity. Instead of adding location state, the ingest writes one item per city (live#weather-cebu.txt) with the city name repeated in the body. Someone texting "ulan ba sa Cebu?" retrieves the Cebu item because its embedding is closest by cosine similarity to a question containing the place name. Cities live in the CITIES table in src/ingest.py — 32 by default: all 17 NCR LGUs (16 cities plus Pateros) and 15 provincial centers across Luzon, Visayas, and Mindanao. Add coordinates there to extend it.
A caveat on NCR granularity. Metro Manila spans roughly 25 km, close to Open-Meteo's coarsest grid cell, so adjacent LGUs will frequently return identical readings — Makati and Mandaluyong are not going to differ. That's the weather model's resolution, not a defect in the ingest. All 17 are enumerated for recognition rather than precision: someone in Navotas gets an answer that says "Navotas," which matters more for trust than a decimal place of rainfall. To skip that, set
WEATHER_CITIEStoncrplus your provincial keys and one Metro Manila item covers the region.
17b. Create the IAM role. IAM → Roles → Create role → AWS service → Lambda → attach AWSLambdaBasicExecutionRole → name it pare-ingest-role. Then add an inline policy (pare-ingest-policy), substituting your values:
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "bedrock:InvokeModel",
"Resource": "arn:aws:bedrock:<REGION>::foundation-model/cohere.embed-multilingual-v3"
},
{
"Effect": "Allow",
"Action": "dynamodb:PutItem",
"Resource": "arn:aws:dynamodb:<REGION>:<ACCOUNT_ID>:table/pare-vectors"
}
]
}No generation-model permission here — this function only calls the embedding model, never Nova Lite.
17c. Create the Lambda. Same flow as §8: Create function → pare-ingest, Python 3.12, arm64, existing role pare-ingest-role. In the editor create ingest.py, paste the contents of src/ingest.py, delete lambda_function.py, Deploy. Then:
- Runtime settings → Handler:
ingest.lambda_handler - Configuration → General configuration → Timeout:
5min (feed fetches; there's no API Gateway ceiling here)
17d. Environment variables (Configuration → Environment variables):
| Key | Value |
|---|---|
VECTOR_TABLE |
pare-vectors |
NEWS_FEED_URL |
https://newsinfo.inquirer.net/feed |
WEATHER_CITIES |
omit the row for all 32 cities, or set e.g. ncr,cebu,davao to narrow |
PAGASA_ADVISORY_URL |
omit the row to use the built-in default; set it to blank to disable |
Only the first row is required. Both weather sources work with no configuration at all — omit the last three rows entirely and you still get per-city forecasts plus typhoon signals, just no news.
WEATHER_CITIESis the one switch where unset and empty differ: omitting the row enables every city, while setting it to an empty string disables weather entirely.
pare-sms-webhook too.
17e. Test it. Test tab → create an event with body {} → Test. A successful run returns something like {"written": ["live#typhoon-signals.txt", "live#weather-ncr.txt", ...], "failed": 0}. Then verify the whole chain:
- DynamoDB console →
pare-vectorstable → Explore table items → findlive#weather-ncr.txt→ open it. Itstextattribute should read roughly:and itsWeather for Metro Manila, Philippines. Current as of 2026-08-01 14:00 PHT. Now in Metro Manila: Thunderstorm, 30.5 C. Today in Metro Manila: Thunderstorm, 8.4 mm rain expected. Tomorrow in Metro Manila: Moderate rain, 10.6 mm rain expected.embeddingattribute should be a long list of numbers. - Text (or use §10's test event against)
pare-sms-webhookwith "ulan ba sa Cebu bukas?". You should get a grounded answer instead of a refusal, and it should be about Cebu specifically — if it answers with another city's weather, the place names aren't discriminating well and you may need fewer cities or more distinctive item text.
17f. Put it on a schedule. Console → Amazon EventBridge → Schedules → Create schedule:
- Name:
pare-ingest-hourly - Schedule pattern: Recurring, Rate-based, every
1hour - Target: AWS Lambda → Invoke → function
pare-ingest - Payload:
{} - Action after completion: NONE. Retry policy and permissions: defaults are fine (the wizard creates a role that can invoke the function).
- Create schedule
Tier the cadence to how fast the data actually moves — hourly for advisories (every 15 minutes when a signal is up), a few times a day for news, weekly for prices. Create a separate schedule per cadence if you add sources with very different volatility.
17g. Optional alarm. Add a CloudWatch alarm on the pare-ingest function's Errors metric (Sum, period 1 hour, threshold > 0, 2 of 2 datapoints, missing data not breaching), named pare-ingest-errors. The function only raises when every configured feed fails, so this firing means answers are silently going stale — the SMS bot keeps working, it just starts refusing time-sensitive questions again.
What to watch after this is live. Source formats change without notice, and the scraped PAGASA page is the most fragile of the three — a template change on their site turns the advisory item empty. The source_failed and source_empty entries in /aws/lambda/pare-ingest logs are your early warning; a source that starts returning empty content writes nothing and does not raise, so it fails quietly by design. Check the logs occasionally, or extend the alarm to cover it.
On rate limits. With all 32 cities on an hourly schedule this makes ~770 Open-Meteo calls a day, against a free non-commercial allowance of 10,000/day — still comfortable, but a 15-minute schedule would put you at ~3,100 and further growth would start to matter. Narrow WEATHER_CITIES before increasing frequency.
This is what static content questions answer from — the "closed-book" half of the product, distinct from §17's time-sensitive feeds. There's no console step for this at all; it's a local script you run whenever a document changes.
- Put your
.txt/.mdsource documents in a local directory —data/in this repo is a reasonable default, but any directory works. - Make sure your local AWS credentials (the same chain the AWS CLI uses — profile, env vars, or SSO) have
bedrock:InvokeModelon the Cohere embed model anddynamodb:PutItem/GetItem/DeleteItemonpare-vectors. The simplest way is to attach the same permissions from §17b's policy (swapPutItemfor the three DynamoDB actions above) to whatever IAM user/role your local credentials resolve to. - Run:
.venv\Scripts\python scripts\ingest_corpus.py data - The script prints one line per file with its chunk count, then a total. Rerun it any time you add, edit, or remove documents — same-named files overwrite their previous chunks in place, and a file that shrinks has its now-orphaned trailing chunks deleted automatically.
- Verify: DynamoDB console →
pare-vectors→ Explore table items → look for keys starting withcorpus#.