ZeroToken Router is a zero-token-first general-purpose AI agent for Track 1 of the AMD Developer Hackathon: ACT II. It answers suitable tasks with a bundled local Qwen model and escalates only the categories that need a stronger Fireworks model.
The scorer contract is intentionally small:
- read
/input/tasks.jsonon startup; - write
/output/results.jsonbefore exiting; - use only model IDs supplied in
ALLOWED_MODELS; - route every Fireworks call through
FIREWORKS_BASE_URL; - remain inside 4 GB RAM, 2 vCPU, 10 minutes, and a 10 GB compressed image.
flowchart LR
I["/input/tasks.json"] --> C["Deterministic 8-category classifier"]
C --> P{"Routing profile"}
P -->|local| Q["Qwen2.5 3B Q4_K_M via llama.cpp"]
P -->|general escalation| M["MiniMax M3 through injected proxy"]
P -->|code escalation| K["Kimi K2.7 Code through injected proxy"]
Q --> O["Ordered, atomic /output/results.json"]
M --> O
K --> O
The classifier is local and deterministic; routing itself spends no Fireworks tokens. Remote tasks run with a maximum of three concurrent requests while the single local model runs sequentially. A remote failure produces a schema-safe deterministic fallback so a failed endpoint cannot omit a task.
| Profile | Behaviour | Use |
|---|---|---|
safe |
MiniMax for non-code, Kimi for code | Accuracy-first submission and API validation |
hybrid |
Local sentiment and simple math; format-sensitive language tasks and harder reasoning escalate | Primary competition profile |
local |
All eight categories use bundled Qwen | Experimental zero-Fireworks-token profile |
hybrid is the image default. Set ROUTING_PROFILE only for controlled experiments.
The grading harness supplies:
FIREWORKS_API_KEY
FIREWORKS_BASE_URL
ALLOWED_MODELS
The agent never logs the key and refuses to call a model missing from ALLOWED_MODELS. For local development copy .env.example to .env, but never commit the resulting file.
Python 3.11+:
python -m venv .venv
source .venv/bin/activate
pip install -r requirements-dev.txt
pytestPowerShell activation:
python -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install -r requirements-dev.txt
pytestDownload the pinned 2.1 GB GGUF for local-model evaluation:
python scripts/download_model.pyThe model URL is pinned to Hugging Face revision 7dabda4d13d513e3e842b20f0d435c732f172cbe
and verified against SHA-256 626b4a6678b86442240e33df819e00132d3ba7dddfe1cdc4fbb18e0a9615c62d.
Start the local OpenAI-compatible mock:
python tools/mock_fireworks_server.py --port 8089In a second terminal:
export FIREWORKS_API_KEY=mock-key
export FIREWORKS_BASE_URL=http://127.0.0.1:8089/inference/v1
export ALLOWED_MODELS=accounts/fireworks/models/minimax-m3,accounts/fireworks/models/kimi-k2p7-code
export ROUTING_PROFILE=safe
export INPUT_PATH=data/practice_tasks.json
export OUTPUT_PATH=output/results.json
python -m tokenrouter.agentUse $env:NAME="value" for the equivalent PowerShell environment assignments.
The Dockerfile downloads and verifies the official Qwen GGUF during the build and installs a pinned CPU-only llama.cpp wheel.
docker build --platform linux/amd64 -t zero-token-router:hybrid .
mkdir -p output
docker run --rm --memory=4g --cpus=2 \
-v "$PWD/data/practice_tasks.json:/input/tasks.json:ro" \
-v "$PWD/output:/output" \
-e FIREWORKS_API_KEY \
-e FIREWORKS_BASE_URL \
-e ALLOWED_MODELS \
zero-token-router:hybridDo not bundle .env or a personal API key in the image. GitHub Actions publishes only linux/amd64 to GHCR and disables provenance so the registry tag resolves directly to the required platform manifest.
The repository contains the eight official practice shapes, the ten retired public validation tasks, and a 19-task regression set:
python -m tokenrouter.agent
python scripts/evaluate.py \
--tasks data/simulated_eval.json \
--results output/results.json \
--required 17This local evaluator is a regression gate, not a substitute for the official LLM judge. Promote a category from Fireworks to local only after it reaches at least 90% on held-out variants with no formatting failures.
Measured on 2026-07-12 with the allowed Fireworks models, the optimized hybrid profile passed all 10 retired public validation tasks and scored 19/19 on three consecutive simulated evaluation runs. The constrained Docker run completed the 19-task batch in 11.06 seconds with --memory=4g --cpus=2. These are local regression results, not claims about the hidden judging set.
- Push the repository publicly to GitHub.
- Run Publish linux-amd64 image in GitHub Actions.
- Make the resulting GHCR package public.
- Confirm an anonymous
docker pullsucceeds. - Submit the immutable tag or digest, not an untracked local image.
Submission support files are in docs/: a five-slide outline, demo script, and an optional static GitHub Pages site. A paid live inference endpoint is not required.
- Project code: MIT, see LICENSE.
- Qwen2.5-3B-Instruct-GGUF: Apache-2.0; weights are downloaded from the official Qwen repository.
- llama.cpp and llama-cpp-python retain their upstream licenses.