Skip to content

Latest commit

 

History

History
137 lines (100 loc) · 3.09 KB

File metadata and controls

137 lines (100 loc) · 3.09 KB

API Specification — v1

Base URL (local): http://localhost:8080

Authentication: cache and stats endpoints require Authorization: Bearer <api_key>, obtained from /v1/auth/signup.


POST /v1/auth/signup

Create a user and receive an API key. Free tier, no credit card.

Request

{ "email": "me@example.com" }   // email optional

201 Response

{ "user_id": "uuid", "api_key": "vc_...", "plan": "free" }

POST /v1/auth/login

Exchange an API key for a short-lived JWT (dashboard sessions).

Request { "api_key": "vc_..." } 200 Response { "token": "<jwt>", "token_type": "Bearer", "expires_in": 3600 }


POST /v1/cache/store

Cache an output and its embedding.

Request

{
  "output": "The capital of France is Paris.",
  "metadata": {
    "model": "gpt-4",
    "tokens_used": 42,
    "cost_usd": 0.004,
    "latency_ms": 800
  }
}

201 Response

{ "semantic_key": "<sha256>", "id": "<uuid>", "status": "stored" }

cost_usd and latency_ms are echoed back as the saved amounts on a future cache hit, so populate them with the real inference cost/latency.


GET /v1/cache/lookup

Find the nearest cached output by cosine similarity.

Query params (or the same keys in a JSON body for large outputs)

  • output (required) — the new text to match against.
  • similarity_threshold (optional, default 0.85) — min cosine similarity for a hit.

200 Response — hit

{
  "hit": true,
  "cached_output": "The capital of France is Paris.",
  "similarity_score": 0.93,
  "cost_saved_usd": 0.004,
  "latency_saved_ms": 800,
  "embedding_used": "text-embedding-3-small"
}

200 Response — miss

{ "hit": false, "embedding_used": "text-embedding-3-small" }

On a miss, the caller runs the model and then calls /v1/cache/store.


GET /v1/stats

Query params: period = day | week | month (default month).

200 Response

{
  "period": "month",
  "total_api_calls": 1200,
  "lookup_calls": 1000,
  "store_calls": 200,
  "cache_hits": 640,
  "cache_hit_rate": 0.64,
  "total_cost_saved_usd": 2.56,
  "total_latency_saved_ms": 512000,
  "top_cached_queries": [{ "output": "The capital of France...", "hits": 88 }]
}

Errors

Status Meaning
400 Missing/invalid field (e.g. empty output)
401 Missing or invalid API key
409 Email already registered
429 Free-tier daily limit reached
500 Internal error

PowerShell examples

$base = "http://localhost:8080"
$key  = (Invoke-RestMethod -Method Post "$base/v1/auth/signup" -ContentType application/json -Body '{"email":"me@example.com"}').api_key
$h    = @{ Authorization = "Bearer $key" }

Invoke-RestMethod -Method Post "$base/v1/cache/store" -Headers $h -ContentType application/json `
  -Body '{"output":"The capital of France is Paris.","metadata":{"model":"gpt-4","cost_usd":0.004,"latency_ms":800}}'

Invoke-RestMethod "$base/v1/cache/lookup?output=What is the capital of France?" -Headers $h

Invoke-RestMethod "$base/v1/stats?period=day" -Headers $h