If you've built web applications, you're familiar with APIs - you send a request, the API processes it, and returns a response. An AI Agent is similar, but much smarter:
- Traditional API: You call
/api/calculate-refundwith specific parameters → It runs predefined logic → Returns a result - AI Agent: You send natural language like "Can I return my laptop?" → It decides which tools to use → Calls multiple functions → Returns a conversational response
Think of an agent as a smart orchestrator that:
- Understands natural language (like a chatbot)
- Decides which functions/tools to call (like a workflow engine)
- Remembers context (like session storage)
- Accesses external data (like API calls)
This project builds a customer service agent for handling returns and refunds. Instead of building separate API endpoints for each task, you build ONE agent that can:
- Answer policy questions
- Check return eligibility
- Calculate refunds
- Look up orders
- Remember customer preferences
The magic: Customers just talk naturally, and the agent figures out what to do.
What it is: Think of this like a super-smart Lambda function that understands language.
bedrock_model = BedrockModel(model_id=MODEL_ID, temperature=0.3)- Model ID:
claude-sonnet-4-5- This is the AI model (like choosing an EC2 instance type) - Temperature:
0.3- How creative vs. consistent (0 = very consistent, 1 = very creative)
How it works:
- You send text: "Can I return my laptop?"
- Claude understands the intent: "User wants to check return eligibility"
- Claude decides: "I need to call the
check_return_eligibilitytool" - Claude formats the response in natural language
In your web app, you might have:
app.post('/api/check-eligibility', (req, res) => { ... })In this agent, you have:
@tool
def check_return_eligibility(purchase_date: str, category: str) -> dict:
"""Check if an item is eligible for return"""
# Business logic here
return {'eligible': True, 'days_remaining': 15}The @tool decorator tells the AI: "You can call this function when needed"
Example flow:
- User: "I bought a laptop on 2024-01-15, can I return it?"
- Agent thinks: "I need purchase date and category"
- Agent calls:
check_return_eligibility('2024-01-15', 'electronics') - Agent gets:
{'eligible': True, 'days_remaining': 15} - Agent responds: "Yes! You have 15 days left to return it."
AgentCore Memory is like having a database that automatically stores:
- Semantic Memory: Facts ("User bought a defective laptop")
- Preferences: User choices ("Prefers email notifications")
- Summary: Conversation context ("Discussing return for order ORD-001")
agentcore_memory_config = AgentCoreMemoryConfig(
memory_id=memory_id,
session_id=session_id,
actor_id=actor_id, # Like a user_id
retrieval_config={
f"app/{actor_id}/semantic": RetrievalConfig(top_k=3),
f"app/{actor_id}/preferences": RetrievalConfig(top_k=3),
}
)Real-world example:
- First conversation: "Hi, I prefer email updates" → Stored in preferences
- Second conversation (days later): "I need help with a return"
- Agent remembers: "I'll send you email updates as you prefer!"
Compare to web dev:
- Traditional: You'd store this in a database and manually query it
- With Memory: It's automatic - the agent retrieves relevant memories when needed
AgentCore Gateway is like AWS API Gateway, but specifically for AI agents. It:
- Secures access with OAuth (Cognito)
- Exposes Lambda functions as "tools" the agent can call
- Uses MCP (Model Context Protocol) - think of it as GraphQL for AI
Example setup:
# Lambda function that looks up orders
def lambda_handler(event, context):
order_id = event['order_id']
# Query database
return {
'product': 'Dell Laptop',
'price': 1299.99,
'purchase_date': '2024-01-15'
}Gateway exposes this as:
{
"name": "lookup_order",
"description": "Look up order details by order ID",
"inputSchema": {
"type": "object",
"properties": {
"order_id": {"type": "string"}
}
}
}Agent can now call it:
@tool
def lookup_order(order_id: str) -> dict:
# Calls Lambda via Gateway
result = gateway_client.call_tool('lookup_order', {'order_id': order_id})
return resultWhat it is: A searchable database of documents using semantic search.
Traditional search:
- User searches: "return policy"
- You search for exact keyword matches
Semantic search:
- User asks: "Can I return opened items?"
- System understands: "User wants return policy for opened products"
- Finds relevant sections even if exact words don't match
@tool
def retrieve(knowledgeBaseId: str, text: str, region: str):
"""Search knowledge base for relevant information"""
# AWS Bedrock does semantic search
# Returns most relevant policy sectionsAgentCore Runtime is like deploying to Lambda, but for AI agents:
@app.entrypoint
def invoke(payload, context=None):
"""This function runs when someone calls your agent"""
user_input = payload.get("prompt")
response = agent(user_input)
return {"result": response}What happens:
- You deploy your agent code
- AWS creates a Docker container
- Stores it in ECR (Elastic Container Registry)
- Runs it serverlessly (auto-scales, pay per request)
- Provides an HTTPS endpoint
Compare to web deployment:
- Traditional: Deploy Express/Flask app to EC2/Lambda
- AgentCore: Deploy agent to Runtime (handles scaling, monitoring, etc.)
What it does: Simple agent with custom tools and knowledge base access.
# Define custom business logic
@tool
def check_return_eligibility(purchase_date: str, category: str) -> dict:
# Calculate if return window is still open
days_since_purchase = (datetime.now() - purchase).days
days_remaining = 30 - days_since_purchase
if days_remaining > 0:
return {'eligible': True, 'days_remaining': days_remaining}
else:
return {'eligible': False, 'reason': 'Return window expired'}Key concept: The @tool decorator makes this function available to the AI. The AI reads the docstring and function signature to understand when and how to use it.
# Create the agent
agent = Agent(
model=bedrock_model, # The AI brain
tools=[retrieve, current_time, check_return_eligibility, ...], # Available functions
system_prompt=system_prompt # Instructions for the AI
)
# Use the agent
response = agent("Can I return my laptop bought on 2024-01-15?")What happens internally:
- Agent receives: "Can I return my laptop bought on 2024-01-15?"
- Agent thinks: "I need to check eligibility. I have a tool for that!"
- Agent calls:
check_return_eligibility('2024-01-15', 'electronics') - Agent gets:
{'eligible': True, 'days_remaining': 15} - Agent responds: "Yes, you can return it! You have 15 days remaining."
What it adds: Memory + Gateway integration
# Memory configuration (like session management)
agentcore_memory_config = AgentCoreMemoryConfig(
memory_id=memory_id,
session_id=session_id, # Like a session cookie
actor_id=actor_id, # Like a user_id
retrieval_config={
f"app/{actor_id}/semantic": RetrievalConfig(top_k=3),
f"app/{actor_id}/preferences": RetrievalConfig(top_k=3),
}
)Memory in action:
# First conversation
user: "Hi, I prefer email notifications"
agent: "Got it! I'll remember that."
# Memory stores: {"preference": "email", "actor_id": "user_001"}
# Second conversation (days later)
user: "I need help with a return"
agent: "I'll help you! I'll send updates via email as you prefer."
# Memory retrieved the preference automatically!Gateway integration (calling external Lambda):
@tool
def lookup_order(order_id: str) -> dict:
"""Look up order via Lambda function through Gateway"""
# Get OAuth token from Cognito
token = get_access_token()
# Call Gateway (which calls Lambda)
result = gateway_client.call_tool('lookup_order', {'order_id': order_id})
return resultFlow:
- User: "Look up order ORD-001"
- Agent calls:
lookup_order('ORD-001') - Gateway authenticates with Cognito
- Gateway invokes Lambda function
- Lambda queries database
- Returns:
{'product': 'Laptop', 'price': 1299.99} - Agent responds: "Your order is for a Laptop at $1,299.99"
What it adds: Production-ready deployment with error handling
@app.entrypoint
def invoke(payload, context=None):
"""Production entrypoint with comprehensive error handling"""
try:
# Extract user input
user_input = payload.get("prompt")
session_id = context.session_id
actor_id = payload.get("actor_id")
# Configure memory
session_manager = AgentCoreMemorySessionManager(...)
# Create agent with all tools
agent = Agent(
model=bedrock_model,
tools=custom_tools + gateway_tools,
system_prompt=system_prompt,
session_manager=session_manager
)
# Process request
response = agent(user_input)
return {"result": response}
except Exception as e:
# Log errors for debugging
print(f"[ERROR] {str(e)}")
return {"error": str(e)}Production features:
- Error handling: Catches failures gracefully
- Logging: Tracks all operations for debugging
- Environment variables: Loads config from environment
- Fallback logic: Works even if Gateway is unavailable
Let's trace a real request through the system:
Step 1: Request arrives at Runtime
payload = {
"prompt": "Can I return my order ORD-001?",
"actor_id": "user_001"
}Step 2: Agent retrieves memories
# Automatically queries Memory
memories = [
{"type": "preference", "content": "Prefers email notifications"},
{"type": "semantic", "content": "Previously returned defective laptop"},
{"type": "summary", "content": "Customer is familiar with return process"}
]Step 3: Agent analyzes request
Claude thinks:
- User wants to return order ORD-001
- I need order details first
- I have a lookup_order tool
- I should call it
Step 4: Agent calls Gateway tool
# Agent executes
result = lookup_order('ORD-001')
# Gateway flow:
# 1. Get OAuth token from Cognito
# 2. Call Lambda function
# 3. Lambda queries database
# 4. Returns order details
result = {
'product': 'Dell XPS 15 Laptop',
'price': 1299.99,
'purchase_date': '2024-03-05',
'category': 'electronics'
}Step 5: Agent checks eligibility
# Agent executes
eligibility = check_return_eligibility('2024-03-05', 'electronics')
eligibility = {
'eligible': True,
'reason': 'Within 30-day return window',
'days_remaining': 15
}Step 6: Agent searches Knowledge Base
# Agent executes
policy = retrieve(
knowledgeBaseId=KB_ID,
text="return policy for electronics",
region="us-west-2"
)
policy = "Electronics can be returned within 30 days..."Step 7: Agent composes response
# Agent combines all information
response = """
Good news! Your Dell XPS 15 Laptop (Order ORD-001) is eligible for return.
Order Details:
- Product: Dell XPS 15 Laptop
- Price: $1,299.99
- Purchase Date: March 5, 2024
- Days Remaining: 15 days
According to our return policy, electronics can be returned within 30 days
of purchase. Your item qualifies for a full refund.
I'll send you the return instructions via email, as you prefer!
"""Step 8: Response sent back
return {"result": response}| Traditional Web App | AI Agent Equivalent | AWS Service |
|---|---|---|
| Express/Flask API | Agent with Tools | Bedrock (Claude) |
| API Routes | @tool functions | Python decorators |
| Session Storage | Memory | AgentCore Memory |
| Database Queries | Knowledge Base | Bedrock KB (Vector DB) |
| API Gateway | Gateway | AgentCore Gateway |
| Lambda Functions | External Tools | Lambda + Gateway |
| OAuth/Cognito | Authentication | Cognito |
| EC2/Lambda Deploy | Runtime Deploy | AgentCore Runtime |
| CloudWatch Logs | Observability | CloudWatch |
system_prompt = """
You are a returns assistant.
You have access to:
1. lookup_order - Get order details
2. check_return_eligibility - Check if returnable
3. calculate_refund_amount - Calculate refund
4. retrieve - Search policy documents
Always be friendly and accurate.
"""This is like giving instructions to a new employee. The AI reads this and understands its role and capabilities.
@tool
def calculate_refund_amount(original_price: float, condition: str, return_reason: str) -> dict:
"""Calculate refund amount based on price, condition, and reason"""
# Business logic
if return_reason == 'defective':
refund_rate = 100 # Full refund
elif return_reason == 'changed_mind':
refund_rate = 85 # 15% restocking fee
refund = original_price * (refund_rate / 100)
return {
'refund_amount': refund,
'refund_percentage': refund_rate,
'explanation': f'{refund_rate}% refund'
}Why this is powerful:
- You write business logic once
- AI decides when to use it
- No need to build separate API endpoints
- Easy to add new tools
retrieval_config={
f"app/{actor_id}/semantic": RetrievalConfig(top_k=3), # Facts
f"app/{actor_id}/preferences": RetrievalConfig(top_k=3), # Preferences
f"app/{actor_id}/{session_id}/summary": RetrievalConfig(top_k=2), # Context
}Think of it like:
-- Semantic table
SELECT * FROM memories
WHERE actor_id = 'user_001'
AND type = 'semantic'
ORDER BY relevance
LIMIT 3;
-- Preferences table
SELECT * FROM memories
WHERE actor_id = 'user_001'
AND type = 'preferences'
ORDER BY relevance
LIMIT 3;Traditional REST API:
POST /api/lookup-order
{
"order_id": "ORD-001"
}
MCP format:
{
"jsonrpc": "2.0",
"method": "tools/call",
"params": {
"name": "lookup_order",
"arguments": {"order_id": "ORD-001"}
}
}Why MCP?
- Standardized way for AI to call tools
- Self-describing (includes schemas)
- Works across different AI models
- Write code
- Build Docker image
- Push to ECR
- Deploy to ECS/Lambda
- Configure API Gateway
- Set up monitoring
# 1. Configure runtime
runtime.configure(
entrypoint="17_runtime_agent.py",
agent_name="returns_agent",
execution_role=role_arn
)
# 2. Deploy (does everything automatically)
runtime.launch(env_vars={
"MEMORY_ID": memory_id,
"GATEWAY_URL": gateway_url,
"COGNITO_CLIENT_ID": client_id
})
# 3. Check status
status = runtime.status() # CREATING → READY
# 4. Invoke
response = runtime.invoke({
"prompt": "Help me with a return",
"actor_id": "user_001"
})What happens behind the scenes:
- Creates CodeBuild project
- Builds Docker container with your code
- Pushes to ECR
- Deploys to AgentCore Runtime
- Sets up auto-scaling
- Configures CloudWatch logging
- Provides HTTPS endpoint
- EC2/ECS: $50-200/month (always running)
- API Gateway: $3.50 per million requests
- Database: $15-100/month
- Total: ~$70-300/month base cost
- Bedrock (Claude): $3 per million input tokens, $15 per million output tokens
- AgentCore Memory: $0.10 per 1000 retrievals
- AgentCore Gateway: $0.01 per 1000 requests
- AgentCore Runtime: $0.0001 per second of compute
- Total: Pay only for what you use (could be $10-50/month for moderate traffic)
- ✅ Users interact with natural language
- ✅ Complex decision-making needed
- ✅ Need to remember context across sessions
- ✅ Multiple tools/services to orchestrate
- ✅ Requirements change frequently
- ✅ Simple, predictable operations
- ✅ Need millisecond response times
- ✅ Exact output format required
- ✅ Cost-sensitive at scale
- ✅ No natural language needed
You built a conversational AI system that:
- Understands natural language (Bedrock/Claude)
- Executes business logic (Custom tools)
- Remembers customers (AgentCore Memory)
- Calls external services (Gateway + Lambda)
- Searches documents (Knowledge Base)
- Deploys serverlessly (AgentCore Runtime)
- Monitors performance (CloudWatch)
In web dev terms: You built a smart API that understands what users want, decides which microservices to call, remembers context, and responds conversationally - all without writing routing logic or state management code.
The AI handles the orchestration; you just provide the tools and business logic!