A minimal, beginner-friendly project that shows off Langfuse's four core capabilities, one script at a time:
| Step | File | What it shows |
|---|---|---|
| 1 | step1_first_trace.py |
Your first trace — one function, one Bedrock call |
| 2 | step2_detailed_tracing.py |
Nested spans, sessions, users, generation detail |
| 3 | step3_evaluation.py |
Scoring traces: a deterministic check + an LLM-as-judge check |
| 4 | step4_prompt_management.py |
Versioned prompts you can edit without touching code |
| 5 | step5_datasets_experiments.py |
Running a fixed test set through your agent as a regression check |
No LangChain, no framework — just plain Python and boto3, so the Langfuse
concepts stay front and center instead of being buried under someone else's
abstraction.
- Python 3.10+
- An AWS account with access to Amazon Bedrock, and model access enabled for Claude 3.5 Haiku (or another Claude model) — check this in the Bedrock console under Model access. Without this step, every call will fail with an access-denied error.
- AWS credentials available locally, either via
aws configure, AWS SSO, or environment variables. This project usesboto3's default credential chain, so anything that already works with the AWS CLI will work here. - A free Langfuse Cloud account: sign up at https://cloud.langfuse.com
- Sign up at https://cloud.langfuse.com and create a new project.
- Go to Settings → API Keys and create a new key pair.
- Note which region you signed up in (EU or US) — it changes the host URL.
cd langfuse-poc
python -m venv venv
source venv/bin/activate # on Windows: venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env
# now edit .env and fill in:
# LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY, LANGFUSE_HOST
# AWS_REGION, BEDROCK_MODEL_IDpython step1_first_trace.py
python step2_detailed_tracing.py
python step3_evaluation.py
python step4_prompt_management.py
python step5_datasets_experiments.pyAfter each script, switch to the Langfuse UI (leave a tab open) and look at what changed. The comment block at the top of each script says exactly where to look and what you should see. Doing it in order matters a little — step 4 assumes step 1-3 already showed you what a generation looks like, and step 5 assumes you're comfortable with the idea of a trace by that point.
- Tracing → Traces: after step 1, open the single trace and expand it. This is the "glass box" view — exact prompt in, exact response out, latency, and (once you've done step 2) token counts and cost.
- Tracing → Sessions and Tracing → Users: after step 2, filter by
demo-session-1anddemo-user. - Scores panel on a trace: after step 3, see both the deterministic and LLM-judge scores attached to the same trace.
- Prompts: after step 4, open
support-reply, look at its version history, and try editing it in the UI, then rerun the script. - Evaluation → Datasets → geography-qa → Runs: after step 5, see the run's pass rate and drill into individual item traces.
- Set up a no-code LLM-as-a-judge evaluator in the Langfuse UI (Evaluation → LLM-as-a-Judge) and point it at the traces from step 1-2, instead of writing your own judge function like in step 3.
- Try the same step 5 dataset with a different
BEDROCK_MODEL_IDand compare the two runs side by side in the Dataset Runs view. - Point this at your actual AgentCore work: swap
bedrock_helper.py's Converse API call for your real agent logic, keep the@observedecorators, and you have production-shaped observability with almost no extra code.
- Every script calls
.flush()before exiting — required for short-lived scripts, since Langfuse batches and sends events asynchronously in the background otherwise. - Bedrock usage in this POC is billed normally by AWS (Claude 3.5 Haiku is inexpensive, well under a cent for all five scripts combined at these prompt sizes). Langfuse Cloud's free tier comfortably covers this volume.
- If
call_bedrockraises an access-denied error, it's almost always the Bedrock model access step in prerequisites, not a Langfuse or code issue.