Skip to content

Latest commit

 

History

History
647 lines (515 loc) · 18.1 KB

File metadata and controls

647 lines (515 loc) · 18.1 KB

AWS AgentCore Workshop - Detailed Explanation for Web Developers

What is an AI Agent? (Think of it like a Smart API)

If you've built web applications, you're familiar with APIs - you send a request, the API processes it, and returns a response. An AI Agent is similar, but much smarter:

  • Traditional API: You call /api/calculate-refund with specific parameters → It runs predefined logic → Returns a result
  • AI Agent: You send natural language like "Can I return my laptop?" → It decides which tools to use → Calls multiple functions → Returns a conversational response

Think of an agent as a smart orchestrator that:

  1. Understands natural language (like a chatbot)
  2. Decides which functions/tools to call (like a workflow engine)
  3. Remembers context (like session storage)
  4. Accesses external data (like API calls)

The Big Picture: What This Project Does

This project builds a customer service agent for handling returns and refunds. Instead of building separate API endpoints for each task, you build ONE agent that can:

  • Answer policy questions
  • Check return eligibility
  • Calculate refunds
  • Look up orders
  • Remember customer preferences

The magic: Customers just talk naturally, and the agent figures out what to do.


Architecture Breakdown (AWS Services You Know)

1. The Brain: Amazon Bedrock (Claude AI Model)

What it is: Think of this like a super-smart Lambda function that understands language.

bedrock_model = BedrockModel(model_id=MODEL_ID, temperature=0.3)
  • Model ID: claude-sonnet-4-5 - This is the AI model (like choosing an EC2 instance type)
  • Temperature: 0.3 - How creative vs. consistent (0 = very consistent, 1 = very creative)

How it works:

  • You send text: "Can I return my laptop?"
  • Claude understands the intent: "User wants to check return eligibility"
  • Claude decides: "I need to call the check_return_eligibility tool"
  • Claude formats the response in natural language

2. The Tools: Python Functions (Like API Endpoints)

In your web app, you might have:

app.post('/api/check-eligibility', (req, res) => { ... })

In this agent, you have:

@tool
def check_return_eligibility(purchase_date: str, category: str) -> dict:
    """Check if an item is eligible for return"""
    # Business logic here
    return {'eligible': True, 'days_remaining': 15}

The @tool decorator tells the AI: "You can call this function when needed"

Example flow:

  1. User: "I bought a laptop on 2024-01-15, can I return it?"
  2. Agent thinks: "I need purchase date and category"
  3. Agent calls: check_return_eligibility('2024-01-15', 'electronics')
  4. Agent gets: {'eligible': True, 'days_remaining': 15}
  5. Agent responds: "Yes! You have 15 days left to return it."

3. Memory: DynamoDB-like Storage for Conversations

AgentCore Memory is like having a database that automatically stores:

  • Semantic Memory: Facts ("User bought a defective laptop")
  • Preferences: User choices ("Prefers email notifications")
  • Summary: Conversation context ("Discussing return for order ORD-001")
agentcore_memory_config = AgentCoreMemoryConfig(
    memory_id=memory_id,
    session_id=session_id,
    actor_id=actor_id,  # Like a user_id
    retrieval_config={
        f"app/{actor_id}/semantic": RetrievalConfig(top_k=3),
        f"app/{actor_id}/preferences": RetrievalConfig(top_k=3),
    }
)

Real-world example:

  • First conversation: "Hi, I prefer email updates" → Stored in preferences
  • Second conversation (days later): "I need help with a return"
  • Agent remembers: "I'll send you email updates as you prefer!"

Compare to web dev:

  • Traditional: You'd store this in a database and manually query it
  • With Memory: It's automatic - the agent retrieves relevant memories when needed

4. Gateway: API Gateway for External Services

AgentCore Gateway is like AWS API Gateway, but specifically for AI agents. It:

  • Secures access with OAuth (Cognito)
  • Exposes Lambda functions as "tools" the agent can call
  • Uses MCP (Model Context Protocol) - think of it as GraphQL for AI

Example setup:

# Lambda function that looks up orders
def lambda_handler(event, context):
    order_id = event['order_id']
    # Query database
    return {
        'product': 'Dell Laptop',
        'price': 1299.99,
        'purchase_date': '2024-01-15'
    }

Gateway exposes this as:

{
  "name": "lookup_order",
  "description": "Look up order details by order ID",
  "inputSchema": {
    "type": "object",
    "properties": {
      "order_id": {"type": "string"}
    }
  }
}

Agent can now call it:

@tool
def lookup_order(order_id: str) -> dict:
    # Calls Lambda via Gateway
    result = gateway_client.call_tool('lookup_order', {'order_id': order_id})
    return result

5. Knowledge Base: Vector Search (Like Elasticsearch)

What it is: A searchable database of documents using semantic search.

Traditional search:

  • User searches: "return policy"
  • You search for exact keyword matches

Semantic search:

  • User asks: "Can I return opened items?"
  • System understands: "User wants return policy for opened products"
  • Finds relevant sections even if exact words don't match
@tool
def retrieve(knowledgeBaseId: str, text: str, region: str):
    """Search knowledge base for relevant information"""
    # AWS Bedrock does semantic search
    # Returns most relevant policy sections

6. Runtime: Serverless Deployment (Like Lambda)

AgentCore Runtime is like deploying to Lambda, but for AI agents:

@app.entrypoint
def invoke(payload, context=None):
    """This function runs when someone calls your agent"""
    user_input = payload.get("prompt")
    response = agent(user_input)
    return {"result": response}

What happens:

  1. You deploy your agent code
  2. AWS creates a Docker container
  3. Stores it in ECR (Elastic Container Registry)
  4. Runs it serverlessly (auto-scales, pay per request)
  5. Provides an HTTPS endpoint

Compare to web deployment:

  • Traditional: Deploy Express/Flask app to EC2/Lambda
  • AgentCore: Deploy agent to Runtime (handles scaling, monitoring, etc.)

Code Walkthrough: The Three Agent Versions

Version 1: Basic Agent (01_returns_refunds_agent.py)

What it does: Simple agent with custom tools and knowledge base access.

# Define custom business logic
@tool
def check_return_eligibility(purchase_date: str, category: str) -> dict:
    # Calculate if return window is still open
    days_since_purchase = (datetime.now() - purchase).days
    days_remaining = 30 - days_since_purchase
    
    if days_remaining > 0:
        return {'eligible': True, 'days_remaining': days_remaining}
    else:
        return {'eligible': False, 'reason': 'Return window expired'}

Key concept: The @tool decorator makes this function available to the AI. The AI reads the docstring and function signature to understand when and how to use it.

# Create the agent
agent = Agent(
    model=bedrock_model,  # The AI brain
    tools=[retrieve, current_time, check_return_eligibility, ...],  # Available functions
    system_prompt=system_prompt  # Instructions for the AI
)

# Use the agent
response = agent("Can I return my laptop bought on 2024-01-15?")

What happens internally:

  1. Agent receives: "Can I return my laptop bought on 2024-01-15?"
  2. Agent thinks: "I need to check eligibility. I have a tool for that!"
  3. Agent calls: check_return_eligibility('2024-01-15', 'electronics')
  4. Agent gets: {'eligible': True, 'days_remaining': 15}
  5. Agent responds: "Yes, you can return it! You have 15 days remaining."

Version 2: Full-Featured Agent (14_full_agent.py)

What it adds: Memory + Gateway integration

# Memory configuration (like session management)
agentcore_memory_config = AgentCoreMemoryConfig(
    memory_id=memory_id,
    session_id=session_id,  # Like a session cookie
    actor_id=actor_id,      # Like a user_id
    retrieval_config={
        f"app/{actor_id}/semantic": RetrievalConfig(top_k=3),
        f"app/{actor_id}/preferences": RetrievalConfig(top_k=3),
    }
)

Memory in action:

# First conversation
user: "Hi, I prefer email notifications"
agent: "Got it! I'll remember that."
# Memory stores: {"preference": "email", "actor_id": "user_001"}

# Second conversation (days later)
user: "I need help with a return"
agent: "I'll help you! I'll send updates via email as you prefer."
# Memory retrieved the preference automatically!

Gateway integration (calling external Lambda):

@tool
def lookup_order(order_id: str) -> dict:
    """Look up order via Lambda function through Gateway"""
    # Get OAuth token from Cognito
    token = get_access_token()
    
    # Call Gateway (which calls Lambda)
    result = gateway_client.call_tool('lookup_order', {'order_id': order_id})
    
    return result

Flow:

  1. User: "Look up order ORD-001"
  2. Agent calls: lookup_order('ORD-001')
  3. Gateway authenticates with Cognito
  4. Gateway invokes Lambda function
  5. Lambda queries database
  6. Returns: {'product': 'Laptop', 'price': 1299.99}
  7. Agent responds: "Your order is for a Laptop at $1,299.99"

Version 3: Production Agent (17_runtime_agent.py)

What it adds: Production-ready deployment with error handling

@app.entrypoint
def invoke(payload, context=None):
    """Production entrypoint with comprehensive error handling"""
    try:
        # Extract user input
        user_input = payload.get("prompt")
        session_id = context.session_id
        actor_id = payload.get("actor_id")
        
        # Configure memory
        session_manager = AgentCoreMemorySessionManager(...)
        
        # Create agent with all tools
        agent = Agent(
            model=bedrock_model,
            tools=custom_tools + gateway_tools,
            system_prompt=system_prompt,
            session_manager=session_manager
        )
        
        # Process request
        response = agent(user_input)
        return {"result": response}
        
    except Exception as e:
        # Log errors for debugging
        print(f"[ERROR] {str(e)}")
        return {"error": str(e)}

Production features:

  • Error handling: Catches failures gracefully
  • Logging: Tracks all operations for debugging
  • Environment variables: Loads config from environment
  • Fallback logic: Works even if Gateway is unavailable

The Complete Workflow: From User to Response

Let's trace a real request through the system:

User Request: "Can I return my order ORD-001?"

Step 1: Request arrives at Runtime

payload = {
    "prompt": "Can I return my order ORD-001?",
    "actor_id": "user_001"
}

Step 2: Agent retrieves memories

# Automatically queries Memory
memories = [
    {"type": "preference", "content": "Prefers email notifications"},
    {"type": "semantic", "content": "Previously returned defective laptop"},
    {"type": "summary", "content": "Customer is familiar with return process"}
]

Step 3: Agent analyzes request

Claude thinks:
- User wants to return order ORD-001
- I need order details first
- I have a lookup_order tool
- I should call it

Step 4: Agent calls Gateway tool

# Agent executes
result = lookup_order('ORD-001')

# Gateway flow:
# 1. Get OAuth token from Cognito
# 2. Call Lambda function
# 3. Lambda queries database
# 4. Returns order details

result = {
    'product': 'Dell XPS 15 Laptop',
    'price': 1299.99,
    'purchase_date': '2024-03-05',
    'category': 'electronics'
}

Step 5: Agent checks eligibility

# Agent executes
eligibility = check_return_eligibility('2024-03-05', 'electronics')

eligibility = {
    'eligible': True,
    'reason': 'Within 30-day return window',
    'days_remaining': 15
}

Step 6: Agent searches Knowledge Base

# Agent executes
policy = retrieve(
    knowledgeBaseId=KB_ID,
    text="return policy for electronics",
    region="us-west-2"
)

policy = "Electronics can be returned within 30 days..."

Step 7: Agent composes response

# Agent combines all information
response = """
Good news! Your Dell XPS 15 Laptop (Order ORD-001) is eligible for return.

Order Details:
- Product: Dell XPS 15 Laptop
- Price: $1,299.99
- Purchase Date: March 5, 2024
- Days Remaining: 15 days

According to our return policy, electronics can be returned within 30 days 
of purchase. Your item qualifies for a full refund.

I'll send you the return instructions via email, as you prefer!
"""

Step 8: Response sent back

return {"result": response}

AWS Services Comparison

Traditional Web App AI Agent Equivalent AWS Service
Express/Flask API Agent with Tools Bedrock (Claude)
API Routes @tool functions Python decorators
Session Storage Memory AgentCore Memory
Database Queries Knowledge Base Bedrock KB (Vector DB)
API Gateway Gateway AgentCore Gateway
Lambda Functions External Tools Lambda + Gateway
OAuth/Cognito Authentication Cognito
EC2/Lambda Deploy Runtime Deploy AgentCore Runtime
CloudWatch Logs Observability CloudWatch

Key Concepts Explained

1. System Prompt (Like API Documentation)

system_prompt = """
You are a returns assistant. 

You have access to:
1. lookup_order - Get order details
2. check_return_eligibility - Check if returnable
3. calculate_refund_amount - Calculate refund
4. retrieve - Search policy documents

Always be friendly and accurate.
"""

This is like giving instructions to a new employee. The AI reads this and understands its role and capabilities.

2. Tools (Like Microservices)

@tool
def calculate_refund_amount(original_price: float, condition: str, return_reason: str) -> dict:
    """Calculate refund amount based on price, condition, and reason"""
    
    # Business logic
    if return_reason == 'defective':
        refund_rate = 100  # Full refund
    elif return_reason == 'changed_mind':
        refund_rate = 85   # 15% restocking fee
    
    refund = original_price * (refund_rate / 100)
    
    return {
        'refund_amount': refund,
        'refund_percentage': refund_rate,
        'explanation': f'{refund_rate}% refund'
    }

Why this is powerful:

  • You write business logic once
  • AI decides when to use it
  • No need to build separate API endpoints
  • Easy to add new tools

3. Memory Namespaces (Like Database Tables)

retrieval_config={
    f"app/{actor_id}/semantic": RetrievalConfig(top_k=3),      # Facts
    f"app/{actor_id}/preferences": RetrievalConfig(top_k=3),   # Preferences
    f"app/{actor_id}/{session_id}/summary": RetrievalConfig(top_k=2),  # Context
}

Think of it like:

-- Semantic table
SELECT * FROM memories 
WHERE actor_id = 'user_001' 
AND type = 'semantic' 
ORDER BY relevance 
LIMIT 3;

-- Preferences table
SELECT * FROM memories 
WHERE actor_id = 'user_001' 
AND type = 'preferences' 
ORDER BY relevance 
LIMIT 3;

4. MCP (Model Context Protocol) (Like GraphQL for AI)

Traditional REST API:

POST /api/lookup-order
{
  "order_id": "ORD-001"
}

MCP format:

{
  "jsonrpc": "2.0",
  "method": "tools/call",
  "params": {
    "name": "lookup_order",
    "arguments": {"order_id": "ORD-001"}
  }
}

Why MCP?

  • Standardized way for AI to call tools
  • Self-describing (includes schemas)
  • Works across different AI models

Deployment Process (Like Deploying a Web App)

Traditional Web App Deployment:

  1. Write code
  2. Build Docker image
  3. Push to ECR
  4. Deploy to ECS/Lambda
  5. Configure API Gateway
  6. Set up monitoring

Agent Deployment:

# 1. Configure runtime
runtime.configure(
    entrypoint="17_runtime_agent.py",
    agent_name="returns_agent",
    execution_role=role_arn
)

# 2. Deploy (does everything automatically)
runtime.launch(env_vars={
    "MEMORY_ID": memory_id,
    "GATEWAY_URL": gateway_url,
    "COGNITO_CLIENT_ID": client_id
})

# 3. Check status
status = runtime.status()  # CREATING → READY

# 4. Invoke
response = runtime.invoke({
    "prompt": "Help me with a return",
    "actor_id": "user_001"
})

What happens behind the scenes:

  1. Creates CodeBuild project
  2. Builds Docker container with your code
  3. Pushes to ECR
  4. Deploys to AgentCore Runtime
  5. Sets up auto-scaling
  6. Configures CloudWatch logging
  7. Provides HTTPS endpoint

Cost Comparison

Traditional API:

  • EC2/ECS: $50-200/month (always running)
  • API Gateway: $3.50 per million requests
  • Database: $15-100/month
  • Total: ~$70-300/month base cost

AI Agent:

  • Bedrock (Claude): $3 per million input tokens, $15 per million output tokens
  • AgentCore Memory: $0.10 per 1000 retrievals
  • AgentCore Gateway: $0.01 per 1000 requests
  • AgentCore Runtime: $0.0001 per second of compute
  • Total: Pay only for what you use (could be $10-50/month for moderate traffic)

When to Use AI Agents vs Traditional APIs

Use AI Agents When:

  • ✅ Users interact with natural language
  • ✅ Complex decision-making needed
  • ✅ Need to remember context across sessions
  • ✅ Multiple tools/services to orchestrate
  • ✅ Requirements change frequently

Use Traditional APIs When:

  • ✅ Simple, predictable operations
  • ✅ Need millisecond response times
  • ✅ Exact output format required
  • ✅ Cost-sensitive at scale
  • ✅ No natural language needed

Summary: What You Built

You built a conversational AI system that:

  1. Understands natural language (Bedrock/Claude)
  2. Executes business logic (Custom tools)
  3. Remembers customers (AgentCore Memory)
  4. Calls external services (Gateway + Lambda)
  5. Searches documents (Knowledge Base)
  6. Deploys serverlessly (AgentCore Runtime)
  7. Monitors performance (CloudWatch)

In web dev terms: You built a smart API that understands what users want, decides which microservices to call, remembers context, and responds conversationally - all without writing routing logic or state management code.

The AI handles the orchestration; you just provide the tools and business logic!