| title | CLI, HTTP Gateway & MCP Reference |
|---|---|
| description | Complete command-line flag reference for dgem (--vertex-url, dgem serve --default-backend vertex_first --cascade-model gemini-3.8-flash, dgem systemone serve), HTTP Gateway proxy routes (/api/decide, /v1/systemone), and Model Context Protocol (MCP) tool parameters. |
This reference documents the global dgem CLI flags, the dgem serve Decision Studio & Gateway server options, the dgem systemone serve tournament adapter, the /v1/systemone and /api/decide HTTP routes, and the Model Context Protocol (MCP) tool schemas.
All dgem subcommands (decide, ask, serve, systemone, mcp, bench, bench-calibration, bench-rerank, bench-bbox, bench-intents, bench-ecotone, bench-jev) share the following global flags and DGEM_* environment variables:
| Flag | Env Var | Default | Description |
|---|---|---|---|
--vertex-url |
DGEM_VERTEX_URL |
"" |
Target Vertex AI Dedicated Endpoint ID (<endpoint-id>) or full /invoke/v1 URL (https://<endpoint-id>.<region>-<project-number>.prediction.vertexai.goog/v1/projects/<project-number>/locations/us-central1/endpoints/<endpoint-id>/invoke/v1). When specified on CLI commands (dgem decide --vertex-url <endpoint-id> --gcp-auth), dgem mints an OAuth2 cloud-platform access token (gcloud auth print-access-token) and routes directly to /invoke/v1/*. |
-u, --url |
DGEM_URL |
http://127.0.0.1:8080/v1 |
Base URL of the upstream server (Local Apple Silicon diffgemma, Serverless Cloud Run dgemma, or dgemma-gateway). |
-m, --model |
DGEM_MODEL |
diffgemma-26b-a4b-it-q4 |
Model identifier (/model or /mnt/gcs/dgemma on Vertex AI / Cloud Run). |
--gcp-auth |
DGEM_GCP_AUTH |
false |
Automatically mint Google Cloud authentication tokens (gcloud auth print-access-token for Vertex AI /invoke/* and Stage 2 Gemini generateContent; gcloud auth print-identity-token for Cloud Run). |
-k, --token |
DGEM_TOKEN |
"" |
Explicit Bearer token override for secured endpoints. |
-s, --stats |
DGEM_STATS |
false |
Print comprehensive GPU timing, KV cache hit rate, restricted-softmax probabilities, and Shannon entropy ( |
--timeout |
DGEM_TIMEOUT |
120s |
HTTP client timeout duration. |
# 1. Single-pass decision on Vertex AI Dedicated Endpoint (0.0s cold start, ~490ms GPU denoise):
./bin/dgem decide --vertex-url <endpoint-id> --gcp-auth \
-t templates/support_triage.json.tmpl \
-v 'ticket=I was billed twice for my annual renewal this morning!' \
--stats
# 2. Full 30-case multi-domain benchmark on Vertex AI Dedicated Endpoint:
./bin/dgem bench --vertex-url <endpoint-id> --gcp-auth \
-d benchmarks/eval_dataset.jsonl \
-o benchmarks/results_vertex_l4_invoke.jsondgem serve starts the unified Web Studio UI, HTTP Gateway REST API, /v1/systemone proxy, and Streamable HTTP MCP server:
./bin/dgem serve \
--default-backend vertex_first \
--vertex-url <endpoint-id> \
--cascade-model gemini-3.8-flash \
-u "https://dgemma-<hash>-uc.a.run.app/v1" \
--gcp-auth \
--port 8080| Flag | Env Var | Default | Description |
|---|---|---|---|
--default-backend |
DGEM_DEFAULT_BACKEND |
vertex_first |
Default upstream GPU routing policy when a request does not specify X-DGem-Backend or backend:• vertex_first (Recommended): Routes to the warm Vertex AI Dedicated Endpoint (--vertex-url) for 0.0 s wakeup and ~490 ms GPU denoise, and automatically falls back to Serverless Cloud Run GPU (-u) if Vertex is updating or scaled to zero.• vertex: Strictly pins requests to the Vertex AI Dedicated Endpoint (/invoke/*).• cloudrun: Strictly pins requests to Serverless Cloud Run GPU (dgemma). |
--vertex-url |
DGEM_VERTEX_URL |
"" |
Target Vertex AI Dedicated Endpoint ID or /invoke/v1 URL used by vertex_first and vertex routing modes. |
--cascade-model |
DGEM_CASCADE_MODEL |
gemini-3.8-flash |
Default Vertex AI Gemini model for Stage 2 Escalation Cascades (gemini-3.8-flash, gemini-3.7-flash, or gemini-3.5-flash-lite). |
--cascade-models |
DGEM_CASCADE_MODELS |
gemini-3.8-flash,gemini-3.7-flash,gemini-3.5-flash-lite |
Comma-separated list of selectable Stage 2 Vertex AI Gemini 3.x models exposed in GET /api/backend-config and the Web Studio. |
--systemone-mode |
DGEM_SYSTEMONE_MODE |
adapter |
Mode for the /v1/systemone route: adapter (bracket tournaments and multi-slot batching, default) or passthrough (direct proxy). |
--systemone-temperature |
DGEM_SYSTEMONE_TEMP |
1.0 |
Post-hoc slot logit temperature scaling /v1/systemone decisions. |
-u, --url |
DGEM_URL |
http://127.0.0.1:8080/v1 |
Upstream Serverless Cloud Run GPU /v1 URL used by cloudrun routing and vertex_first failover. |
--port |
PORT |
8080 |
HTTP listener port. |
dgem systemone serve runs a standalone HTTP adapter that implements POST /v1/systemone (the protocol used by apolinario/decision-index). It splits choices with more than 26 options into 2-stage bracket tournaments and batches more than 8 questions across forward passes, sending sub-requests to an upstream DiffusionGemma server over /v1/chat/completions.
# In front of a local structured server / vLLM:
dgem systemone serve --port 8080 --upstream http://127.0.0.1:8081/v1
# In front of a Vertex AI Dedicated Endpoint (a GCP access token is added automatically):
dgem systemone serve --port 8095 \
--upstream "https://<endpoint-id>.<region>-<project-number>.prediction.vertexai.goog/v1/projects/<project-id>/locations/<region>/endpoints/<endpoint-id>/invoke/v1"
# Require a Bearer key on incoming requests:
dgem systemone serve --port 8080 --api-key "my-secret-key"| Flag | Env Var | Default | Description |
|---|---|---|---|
-p, --port |
PORT, then SYSTEMONE_PORT
|
8080 |
Port to listen on. |
--host |
— | 0.0.0.0 |
Interface to bind. |
--upstream |
SYSTEMONE_UPSTREAM_URL, then UPSTREAM_DGEMMA_URL
|
-u / --url if set, else http://127.0.0.1:8081/v1
|
Upstream /v1 base URL (local server, Cloud Run, or a Vertex .../invoke/v1 URL). |
--temperature |
— | 1.0 |
Post-hoc slot temperature scaling 1.0 = unscaled). |
--max-slots |
— | 8 |
Maximum questions per forward pass before batching. |
--bracket-size |
— | 20 |
Maximum options per round-1 tournament bracket. |
--api-key |
SYSTEMONE_API_KEY, then API_KEY
|
"" (open) |
If set, requests must send Authorization: Bearer <key>. |
--null-prior-debias |
— | false |
Divide out the positional option-A prior (validate on your data first). |
--prior-alpha |
— | 0.50 |
Exponent for null-prior de-biasing. |
--dual-mirror |
— | false |
Also read a reversed option ordering (research diagnostic). |
--naive-limits |
— | false |
Reproduce the naive 26-option / 10-question rejections (benchmark ablation only). |
Every inference endpoint on dgem serve (https://<your-dgem-gateway>) accepts backend selection via HTTP header X-DGem-Backend: vertex_first | vertex | cloudrun, query parameter ?backend=vertex_first, or JSON body field "backend": "vertex_first", and returns the X-DGem-Backend-Used: vertex | cloudrun response header.
| Route | Method | Description |
|---|---|---|
/api/decide & /api/decide/{template} |
POST |
Renders a named or inline (custom_template) .json.tmpl policy with variables, executes Stage 1 DiffusionGemma readout on vertex_first / vertex / cloudrun, and optionally runs the Stage 2 Gemini Cascade (cascade_mode: "off" | "entropy" | "on_miss", cascade_threshold: 0.35, cascade_model: "gemini-3.8-flash"). |
/v1/systemone |
POST |
Direct pass-through proxy to structured_server.py's /v1/systemone (SystemOne / JevBench schema evaluation). Supports both application/json ({"state": ..., "questions": ...}) and multipart/form-data (image file + JSON fields), routing to /invoke/v1/systemone on Vertex AI or /v1/systemone on Cloud Run GPU. |
/v1/chat/completions |
POST |
OpenAI-compatible structured diffusion decision envelope proxy with vertex_first auto-failover and automatic GCP token injection. |
/v1/raw/chat/completions |
POST |
Direct pass-through proxy to vLLM's raw /v1/chat/completions endpoint. |
/api/backend-config (live vertex_status), /api/vertex/deploy, /api/vertex/teardown |
GET / POST |
Live Vertex AI Dedicated Endpoint replica telemetry and 1-click provisioning/teardown (g4-standard-48 1× NVIDIA RTX PRO 6000). |
The MCP inference tools (decide_policy, decide_custom_questions, and locate_bounding_boxes) accept the following backend routing and Stage 2 Gemini Cascade parameters:
| MCP Argument | Type | Allowed Values / Default | Description |
|---|---|---|---|
backend |
string |
"vertex_first" (default) | "vertex" | "cloudrun"
|
Selects the GPU execution target (vertex_first routes to warm Vertex AI Dedicated Endpoint with automatic Cloud Run failover). |
vertex_url |
string |
"" (optional)
|
Custom Vertex AI Dedicated Endpoint ID or /invoke/v1 URL override. |
cascade_mode |
string |
"off" (default) | "entropy" | "on_miss"
|
Stage 2 Gemini Cascade trigger policy ("entropy" escalates when Stage 1 Shannon entropy cascade_threshold). |
cascade_threshold |
number |
0.35 (default, in nats)
|
Shannon entropy threshold "entropy" escalation. |
cascade_model |
string |
"gemini-3.8-flash" (default)
|
Stage 2 Vertex AI Gemini model ("gemini-3.8-flash", "gemini-3.7-flash", or "gemini-3.5-flash-lite"). |
| Variable | Default | Purpose |
|---|---|---|
DGEM_VERTEX_URL |
"" |
Target Vertex AI Dedicated Endpoint ID or /invoke/* URL. |
DGEM_VERTEX_ENDPOINT_ID |
"" |
Dedicated Endpoint ID override. |
DGEM_VERTEX_MODEL_ID |
"" |
Target Vertex AI Model ID deployed by /api/vertex/deploy. |
DGEM_VERTEX_SA |
"" |
Dedicated Service Account for Vertex AI model deployment. |
DGEM_DEFAULT_BACKEND |
vertex_first |
Default backend selector (vertex_first, vertex, cloudrun, local). |
DGEM_CASCADE_MODEL |
gemini-3.8-flash |
Default Gemini model for Stage 2 escalation. |
DGEM_CASCADE_MODELS |
gemini-3.8-flash,gemini-3.7-flash,gemini-3.5-flash-lite |
List of enabled Gemini cascade models. |
DGEM_GATEWAY_HOSTS |
"" |
Comma-separated extra hostnames recognized as remote GCP endpoints for token injection. |
DGEM_REMOTE_URL |
"" |
Remote gateway endpoint used by dgem mcp --remote. |
GCP_PROJECT |
"" |
Google Cloud project ID (falls back to GCP metadata server). |
GCP_PROJECT_NUMBER |
"" |
Google Cloud numeric project ID (detected automatically if on GCP). |
GCP_REGION |
us-central1 |
Default Google Cloud region for services and endpoints. |