Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Cloud Run Ollama Agent Template

Deploy any Ollama model to Google Cloud Run GPU behind a private IAM-protected proxy. The proxy exposes OpenAI-compatible /v1/responses and /v1/chat/completions endpoints for normal HTTP clients, and it also normalizes Responses/tool-call payloads so Codex can use the same service as a custom provider.

V1 is Ollama-only. There is no vLLM, Gemma fallback, legacy path, or public API key mode in this template.

Current Status

  • The earlier deepseek-r1:671b experiment was useful for sizing and timeout discovery, but it did not produce a usable Cloud Run deployment in this repo.
  • The currently validated path is mixtral:8x22b behind the private proxy in europe-west4 on 1x nvidia-rtx-pro-6000.
  • That Mixtral setup has been deployed successfully end-to-end: private runtime, private proxy, OpenAI-compatible /v1/responses and /v1/chat/completions, and a working Codex profile.
  • The validated Mixtral deployment here uses MAX_INSTANCES="1" and CONCURRENCY="1". Treat it as a single-request runtime unless you are ready to tune cost, latency, and capacity tradeoffs.

Operational Notes

These notes are intentionally written for both humans and AI agents reading the repo.

  • Cloud Run in europe-west4 can run with nvidia-rtx-pro-6000.
  • This repo has been validated to accept 30 CPU and 120Gi memory on Cloud Run with 1x nvidia-rtx-pro-6000.
  • 120Gi is accepted by Cloud Run for this GPU type. Do not assume the limit is 96Gi just because the GPU has 96 GB VRAM. GPU VRAM and instance memory are separate.
  • Cloud Run supports exactly one GPU per service instance in this template.
  • For nvidia-rtx-pro-6000, Cloud Run requires at least 20 CPU and 80Gi memory.
  • The runtime deploy script uses --no-gpu-zonal-redundancy. This maps to "No Zonal Redundancy" in the Cloud Run console.
  • STARTUP_INITIAL_DELAY_SECONDS must not exceed 240. Higher values are rejected by Cloud Run at deploy time.
  • Large Ollama models are pulled during Cloud Build, not at runtime. For very large models, the first blocker is often Cloud Build disk and timeout rather than Cloud Run deployment shape.
  • deepseek-r1:671b starts a very large model pull during image build. In our test run, Ollama reported a roughly 404 GB download during ollama pull.
  • Very large model image builds can take many hours. One deepseek-r1:671b build in this repo took about 7 hours 35 minutes before the image tag became available in Artifact Registry.
  • Do not deploy the runtime until the model image build has fully completed and pushed its tag. If Cloud Run deploy runs too early, it can fail with Image ... not found; rerun ./scripts/deploy-runtime.sh after the image is visible in Artifact Registry.
  • A successful Cloud Build does not guarantee that Cloud Run can import an extremely large runtime image within the deployment deadline. In one deepseek-r1:671b attempt, Cloud Run spent about 2 hours importing the image and then failed the revision with Deadline exceeded.
  • For giant baked images, the next bottleneck after Cloud Build can be Cloud Run image import time rather than container startup time. If a revision never reaches startup logs and the status mentions image import progress, consider a smaller model or a different runtime platform.
  • For very large models, use low concurrency. CONCURRENCY="1" is the safest starting point.
  • MODEL_CONTEXT_WINDOW is a client/profile declaration for Codex, not a hard guarantee that the runtime can serve that full context under memory pressure.

Verified Mixtral-8x22b Profile

The following profile is the current known-good end-to-end deployment in this repo for a Cloud Run GPU service plus private proxy:

GCP_REGION="europe-west4"
SERVICE_PREFIX="mixtral-8x22b"
OLLAMA_MODEL="mixtral:8x22b"
MODEL_ALIAS="mixtral-8x22b"
GPU_TYPE="nvidia-rtx-pro-6000"
CLOUD_RUN_CPU="30"
CLOUD_RUN_MEMORY="120Gi"
MAX_INSTANCES="1"
CONCURRENCY="1"
STARTUP_INITIAL_DELAY_SECONDS="240"
PROXY_CPU="2"
PROXY_MEMORY="2Gi"
MODEL_CONTEXT_WINDOW="32768"
REQUEST_TIMEOUT_SECONDS="7200"

Operational notes for this validated Mixtral profile:

  • Use the proxy service as the client entry point. The runtime service is not the public API surface.
  • The model works through the OpenAI-compatible proxy endpoints and Codex can authenticate with gcloud auth print-identity-token.
  • With MAX_INSTANCES="1" and CONCURRENCY="1", back-to-back or parallel requests can hit Cloud Run capacity limits while the single GPU instance is cold-starting or busy. If lower latency matters more than cost, consider setting min-instances=1 on the runtime.

DeepSeek-R1 671B Attempt Notes

The following profile documents the earlier large-model experiment with deepseek-r1:671b on nvidia-rtx-pro-6000. It is included as a sizing note, not as a verified working deployment for this repo:

GCP_REGION="europe-west4"
SERVICE_PREFIX="deepseek-r1"
OLLAMA_MODEL="deepseek-r1:671b"
MODEL_ALIAS="deepseek-r1-671b-iq4xs"
GPU_TYPE="nvidia-rtx-pro-6000"
CLOUD_RUN_CPU="30"
CLOUD_RUN_MEMORY="120Gi"
MAX_INSTANCES="1"
CONCURRENCY="1"
STARTUP_INITIAL_DELAY_SECONDS="240"
MODEL_CONTEXT_WINDOW="4096"
REQUEST_TIMEOUT_SECONDS="7200"

For that profile, the runtime image build should also use a larger Cloud Build worker than the original template defaults. A practical starting point is:

options:
  machineType: E2_HIGHCPU_32
  diskSizeGb: 1000
timeout: 86400s

Quickstart

  1. Sign in to Google Cloud:

    gcloud auth login
    gcloud auth application-default login
  2. Create your config:

    cp config.example.env config.env
  3. Edit config.env.

    Set GCP_PROJECT_ID, choose GCP_REGION, and set any Ollama model you want:

    OLLAMA_MODEL="qwen3-coder:30b"
    MODEL_ALIAS="qwen3-coder-30b"

    You are not required to use gpt-oss. Any Ollama model name is valid as long as the selected Cloud Run CPU, memory, GPU, and region can run it.

    If you want the tested RTX PRO 6000 profile from this repo, start from the "Verified Mixtral-8x22b Profile" above. The DeepSeek section is a record of a failed large-model attempt, not the default recommendation.

  4. Prepare Google Cloud resources:

    ./scripts/bootstrap-gcp.sh
  5. Build the model image:

    ./scripts/build-model-image.sh

    This step waits for Cloud Build to finish, but very large models can still take hours. For example, deepseek-r1:671b took about 7 hours 35 minutes in one run. Wait for this command to complete successfully before deploying the runtime.

  6. Deploy the private Ollama runtime:

    ./scripts/deploy-runtime.sh
  7. Deploy the private proxy:

    ./scripts/deploy-proxy.sh
  8. Add a Codex profile:

    ./scripts/write-codex-profile.sh --install
  9. Smoke test the normal API and Codex:

    ./scripts/smoke-test.sh

Architecture

Codex or HTTP client
        |
        | Cloud Run IAM identity token
        v
Private proxy service
        |
        | service account identity token
        v
Private Ollama runtime service

The runtime is not meant to be called directly. The proxy is the stable API surface for both normal clients and Codex.

Normal API Usage

Fetch an identity token and call the proxy URL:

TOKEN="$(gcloud auth print-identity-token)"
source scripts/lib/config.sh
PROXY_URL="$(gcloud run services describe "$(proxy_service_name)" \
  --project "$GCP_PROJECT_ID" \
  --region "$GCP_REGION" \
  --format 'value(status.url)')"

curl -sS "$PROXY_URL/v1/responses" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model":"mixtral-8x22b","input":"Respond exactly: OK"}'

Chat completions also work:

curl -sS "$PROXY_URL/v1/chat/completions" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model":"mixtral-8x22b","messages":[{"role":"user","content":"Respond exactly: OK"}]}'

Model Selection

OLLAMA_MODEL is the model pulled during Cloud Build. Examples:

  • qwen3-coder:30b
  • mixtral:8x22b
  • gpt-oss:120b
  • llama3.1:8b
  • deepseek-coder:latest

MODEL_ALIAS is what the proxy exposes to clients and Codex. It can match OLLAMA_MODEL, but a colon-free alias is usually quieter in Codex telemetry.

Large models need larger images, longer Cloud Build time, enough GPU VRAM, and larger Cloud Run CPU/memory. Tune CLOUD_RUN_CPU, CLOUD_RUN_MEMORY, GPU_TYPE, MAX_INSTANCES, and CONCURRENCY for your model.

For very large models on Ollama, also tune Cloud Build worker size, build disk, and build timeout. Build-time model download size can be much larger than the final deployment shape suggests.

Security

Services deploy with --no-allow-unauthenticated.

  • Users call the proxy with a Google identity token.
  • The proxy calls the runtime with its service account identity token.
  • The runtime service is not public.

This template does not include public API keys in V1.

References

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages