Run batch jobs for less, without missing deadlines.
AgentGrid prices every way to run a batch job, recommends the cheapest plan that still meets your deadline, submits it to Slurm only with your approval, then checks on the controller what actually ran.
The Run tab after a run: three plans compared, the cheapest one that meets the deadline approved, then submitted and verified on Slurm.
Captured locally against a mock slurmrestd. Costs and times are modelled, as the page says.
Batch jobs are usually sized by habit ("60 cores, to be safe"). The constraints that matter, a deadline and a budget, are not something Slurm or the cloud console reasons about. AgentGrid does the arithmetic before the run and checks the facts after it.
- Cheaper, by construction. Every executable machine shape is priced against your deadline. The recommendation is the cheapest plan that still meets it, and the page says how much less it costs than the fastest plan for the same deadline.
- Nothing runs without you. Plans are proposals. In the default mode a job is submitted only after you approve that exact plan, and a double click or a retried request never starts a second job.
- Checked, not assumed. A job counts as running only when the Slurm controller reports an allocation for it. Modelled numbers are labelled as modelled and never shown as measured.
The dashboard follows the four steps of a run.
| Step | What AgentGrid does |
|---|---|
| 1. Describe | You give the command, the vCPU and memory, how long the work takes on one vCPU, the deadline, the budget, and whether the job can be interrupted. |
| 2. Compare | A deterministic engine scales the work to each machine shape (Amdahl's law, 97% parallel for a parallel job), prices it, drops the shapes that miss the deadline or the budget, and checks the rest against your Compute Engine quota. You get up to three plans: Lowest cost, Fastest and Balanced. When nothing fits, you get what blocks and what to relax, never an invented winner. |
| 3. Approve | You approve one plan. The approval is bound to the plan's content: change the plan and the approval no longer applies. |
| 4. Verify | AgentGrid submits the job through slurmrestd and follows it on the controller: pending, running on the allocated CPUs, completed. The Cluster tab shows the live controller state, the capacity that could take the next job, and why a job may be waiting. |
On Slurm, the machine shapes are your partitions (see Connect your Slurm cluster).
flowchart LR
Op(["Operator"]) --> UI["Dashboard<br/>Run · Ledger · Cluster"]
UI --> API["Control plane<br/>FastAPI"]
API --> Plans["Plan engine<br/>deterministic"]
API --> Gov["Governance and ledger<br/>SQLite"]
API --> RT["Runtime adapter"]
API --> GCP["Compute Engine quota<br/>and Capacity Advisor"]
API -- "/optimize" --> Agent["Gemini agent<br/>Google ADK"]
Agent -- "MCP tools" --> MCP["MCP server"]
MCP --> RT
RT --> Slurm[("Slurm<br/>slurmrestd")]
RT --> Sim[("Built-in simulator")]
- Arithmetic is code. Plans, quota checks, ceilings and verification are deterministic Python. The Gemini
agent is optional (
/optimize) and acts only through typed MCP tools that enforce the same rules. - Three kinds of numbers, never mixed. Measured (read from the runtime), modelled (the plan engine: 0.050 EUR per vCPU-hour by default, not a cloud price) and declared (a figure someone typed). The dashboard and the API label each one.
- Governance held by the server. Advisory never changes anything, Validation (the default) needs your approval for each plan, Delegation runs within a budget, machine types, regions and a retry count. A caller can tighten the mode, never loosen it.
- Exactly-once submission. A submission is claimed in SQLite, keyed by workload and plan fingerprint, before the runtime is touched. A double click, a replayed request or a restart resolve to the same job.
- Ceilings that add up. The delegated budget and the retry count are rebuilt from the submission ledger inside the same transaction as the claim, so concurrent requests cannot slip past them.
- Constraints bind execution. Regions, zones, Spot and fallback rules declared on the workload are checked at submission and on every fallback rung, not only when planning.
How each guarantee is enforced, and the test that pins it: docs/guarantees.md.
| Area | Status |
|---|---|
| Plans, governance, exactly-once submission, ceilings, constraints | Covered by the test suite (379 tests) |
| Compute Engine quota read | Exercised against the live API |
| Slurm adapter | Tested against a mock slurmrestd, in unit tests and in a real browser. Live re-run on the reference cluster: pending |
| Dashboard | Compiled and rendered offline by the test suite, driven in headless Chrome during development |
| Measured cost and duration in the Ledger | Not written back yet: after a verified run, the Ledger still says "none measured yet" |
| History on Cloud Run | SQLite inside the container, lost when the instance is replaced |
| Multi-tenant authentication | Out of scope. An optional shared API key protects the API |
No cloud account needed: without configuration, runs go to the built-in simulator, and the page says so.
git clone https://github.com/samy-fadel/agentgrid.git
cd agentgrid
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
AGENTGRID_PLAN_CAPACITY_CHECK=false PORT=8080 python3 -m compute_agent.appAGENTGRID_PLAN_CAPACITY_CHECK=false skips the Compute Engine quota lookup, which needs Google Cloud credentials
(without them, every comparison waits on the credential search). Plans then say NOT_CHECKED. With credentials,
drop it and set GOOGLE_CLOUD_PROJECT.
Open http://localhost:8080/, or call the API directly:
curl -s localhost:8080/health
curl -s -X POST localhost:8080/api/plans/compare -H 'Content-Type: application/json' \
-d '{"workload_profile": {"workload_id": "wl-demo", "cpu_requested": 8, "memory_mb_requested": 16384,
"estimated_duration_minutes": 60, "deadline_minutes_from_start": 60, "budget_amount": 50}}'An unknown key in workload_profile is rejected with HTTP 400 naming the field, so a typo cannot silently drop
a constraint.
gcloud auth application-default login
cp .env.example .env # set GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION, AGENTIC_COMPUTE_MODEL
adk web . # playground on http://localhost:8000, pick compute_agentThe control plane serves the same agent on POST /optimize and POST /optimize/stream (Server-Sent Events).
AgentGrid talks to Slurm through slurmrestd (data parser v0.0.41), authenticated with a JWT.
COMPUTE_RUNTIME=slurm \
SLURM_REST_URL=http://<controller>:6842/slurm/v0.0.41 \
SLURM_JWT_TOKEN=<output of scontrol token> \
SLURM_USER=<slurm user> \
PORT=8080 python3 -m compute_agent.app- Partitions map to machine types in
SUPPORTED_SLURM_PARTITIONS(slurm_adapter.py):debug→n2-standard-2,compute→c2-standard-60,h3→h3-standard-88. Edit it to match your cluster. - A job is submitted with its partition, CPUs, memory and command. No node features or constraints are sent:
the partition pins the machine type, which is recorded in the job comment. A bare command such as
sleep 30gets a#!/bin/bashline. - Nothing is called at startup, so an unreachable controller does not stop the service:
/api/snapshotreturns the error instead.
A push to main runs the Cloud Build pipeline in cloudbuild.yaml.
flowchart LR
Push["git push main"] --> Tests["run-unit-tests"] --> Build["build and push<br/>two images"] --> Deploy["deploy-services"]
Deploy --> Agent["Cloud Run<br/>agentic-compute-agent<br/>dashboard and API"]
Deploy --> MCP["Cloud Run<br/>agentic-compute-mcp<br/>private"]
Agent -- "run.invoker" --> MCP
Agent --> VPC[("Slurm network<br/>slurmrestd")]
MCP --> VPC
scripts/deploy_services.sh deploys both services with Direct VPC egress to the
cluster network and the JWT from Secret Manager, then grants the agent's service account roles/run.invoker on
the MCP service. One-time setup (Artifact Registry, trigger, IAM): scripts/setup_ci_cd.sh.
| Deploy variable | Default | Effect |
|---|---|---|
AGENT_COMPUTE_RUNTIME |
slurm |
simulator keeps the dashboard on the built-in simulator |
REQUIRE_IAM_AUTH |
false |
true deploys the agent service with --no-allow-unauthenticated (so does ALLOW_UNAUTHENTICATED=false) |
AGENTGRID_API_KEY_SECRET |
none | Secret Manager secret holding an API key that the agent service then requires |
SLURM_REST_URL, SLURM_SECRET_NAME |
in cloudbuild.yaml |
Controller URL, and the secret holding the JWT |
VPC_NETWORK, VPC_SUBNET |
in cloudbuild.yaml |
Network and subnet for Direct VPC egress |
MCP_MIN_INSTANCES |
0 |
Minimum instances of the MCP service |
Important
Unless you set REQUIRE_IAM_AUTH=true, the agent service is deployed with --allow-unauthenticated.
Once it is connected to Slurm, anyone who can reach it can approve and run jobs on your cluster.
| Variable | Default | What it does |
|---|---|---|
COMPUTE_RUNTIME |
simulator |
slurm sends runs to your cluster |
SLURM_REST_URL |
http://10.0.0.4:6842/slurm/v0.0.41 |
slurmrestd base URL |
SLURM_JWT_TOKEN |
empty | JWT, sent as X-SLURM-USER-TOKEN |
SLURM_USER |
slurm |
Slurm user, sent as X-SLURM-USER-NAME |
SLURM_VERIFICATION_TIMEOUT_SECS |
60 |
How long a change may wait for the controller's confirmation before it is reported as timed out |
AGENTGRID_DEFAULT_CONTROL_MODE |
validation |
advisory, validation or delegation |
AGENTGRID_DB_PATH |
~/.agentgrid/agentgrid_history.db |
SQLite file: plans, approvals, submission ledger, history |
AGENTGRID_API_KEY |
none | When set, API calls need X-API-Key: <key> or Authorization: Bearer <key> |
AGENTGRID_PLAN_CAPACITY_CHECK |
true |
Checks each plan against the Compute Engine quota (needs Google Cloud credentials). With false, plans say NOT_CHECKED |
AGENTGRID_QUOTA_CACHE_TTL |
30 |
Seconds a successful quota read is reused (0 disables the cache) |
AGENTGRID_ALLOWED_ORIGINS |
any origin, no credentials | Comma-separated CORS allowlist |
GOOGLE_CLOUD_PROJECT |
PROJECT_ID, else a built-in demo project |
Project for quota reads and Vertex AI. The fallback is reported as built_in_default |
MCP_SERVER_URL |
local stdio | Remote MCP server used by the agent |
AGENTIC_COMPUTE_MODEL |
gemini-2.5-flash |
Model used by the agent |
The dashboard is a client of the same HTTP API. Interactive documentation is served on /docs.
| Endpoint | Purpose |
|---|---|
POST /api/plans/compare |
Compare plans for a workload profile |
POST /api/plans/approve |
Approve one plan, bound to its fingerprint |
POST /api/execute-plan |
Submit an approved plan through the governance gate |
GET /api/snapshot |
Live runtime state: cluster, current job, last action |
POST /api/capacity-search |
Machine types that fit, with quota and Spot signals |
POST /api/diagnose |
Why a job is waiting, and who holds each blocker |
GET, POST /api/workloads/{id}/control |
Read or set the control mode and the delegation policy |
POST /api/workloads/{id}/fallback |
Propose the next authorised fallback plan (never submits) |
GET /api/history |
Past runs, with estimated, measured and declared costs |
POST /optimize, POST /optimize/stream |
Run the Gemini agent (JSON or Server-Sent Events) |
The MCP server (mcp_server.py) exposes the same capabilities as 18
tools, over stdio locally or SSE on Cloud Run.
python3 -m pytest -q # full test suite
python3 tools/check_jsx.py # parses the dashboard's JSX offline
python3 tools/render_check.py # renders every tab offline and reports errorsThe dashboard is a single file, compute_agent/static/index.html: React 18
and Tailwind from CDNs, no build step. The two tools catch syntax errors and render-time crashes without Node.js
or a browser. They do not cover CSS or layout.
compute_agent/
app.py Control plane: dashboard and HTTP API
agent.py Gemini agent (Google ADK) with the MCP toolset
static/index.html Dashboard: Run · Ledger · Cluster
src/agentic_compute/
plan_engine.py Deterministic plan comparison
governance.py Control modes, approvals, submission ledger
execution_controller.py Governance gate, fallback ladder, exactly-once submission
slurm_adapter.py Slurm runtime over slurmrestd, with verification
simulator.py Built-in simulated runtime
capacity_advisor.py Compute Engine quota and Capacity Advisor client
diagnostics.py Blocker diagnosis
history.py SQLite history and cost reconciliation
mcp_server.py MCP server (18 tools)
scripts/ deploy_services.sh, setup_ci_cd.sh
tests/ pytest suite
tools/ Offline dashboard checks
- Deterministic plan comparison, with a quota check on every plan
- Server-side governance, exactly-once submission, cumulative ceilings
- Slurm runtime over
slurmrestd, with verification on the controller - CI/CD with Cloud Build and Cloud Run, with Direct VPC egress
- Write the measured cost and duration back to the Ledger
- Persistent history on Cloud Run
- Kubernetes / GKE runtime adapter
- Multi-agent orchestration: planner, cost guardian, cluster watchdog