Deploy any Ollama model to Google Cloud Run GPU behind a private IAM-protected
proxy. The proxy exposes OpenAI-compatible /v1/responses and
/v1/chat/completions endpoints for normal HTTP clients, and it also normalizes
Responses/tool-call payloads so Codex can use the same service as a custom
provider.
V1 is Ollama-only. There is no vLLM, Gemma fallback, legacy path, or public API key mode in this template.
- The earlier
deepseek-r1:671bexperiment was useful for sizing and timeout discovery, but it did not produce a usable Cloud Run deployment in this repo. - The currently validated path is
mixtral:8x22bbehind the private proxy ineurope-west4on1x nvidia-rtx-pro-6000. - That Mixtral setup has been deployed successfully end-to-end: private runtime,
private proxy, OpenAI-compatible
/v1/responsesand/v1/chat/completions, and a working Codex profile. - The validated Mixtral deployment here uses
MAX_INSTANCES="1"andCONCURRENCY="1". Treat it as a single-request runtime unless you are ready to tune cost, latency, and capacity tradeoffs.
These notes are intentionally written for both humans and AI agents reading the repo.
- Cloud Run in
europe-west4can run withnvidia-rtx-pro-6000. - This repo has been validated to accept
30 CPUand120Gimemory on Cloud Run with1x nvidia-rtx-pro-6000. 120Giis accepted by Cloud Run for this GPU type. Do not assume the limit is96Gijust because the GPU has96 GBVRAM. GPU VRAM and instance memory are separate.- Cloud Run supports exactly one GPU per service instance in this template.
- For
nvidia-rtx-pro-6000, Cloud Run requires at least20 CPUand80Gimemory. - The runtime deploy script uses
--no-gpu-zonal-redundancy. This maps to "No Zonal Redundancy" in the Cloud Run console. STARTUP_INITIAL_DELAY_SECONDSmust not exceed240. Higher values are rejected by Cloud Run at deploy time.- Large Ollama models are pulled during Cloud Build, not at runtime. For very large models, the first blocker is often Cloud Build disk and timeout rather than Cloud Run deployment shape.
deepseek-r1:671bstarts a very large model pull during image build. In our test run, Ollama reported a roughly404 GBdownload duringollama pull.- Very large model image builds can take many hours. One
deepseek-r1:671bbuild in this repo took about 7 hours 35 minutes before the image tag became available in Artifact Registry. - Do not deploy the runtime until the model image build has fully completed and
pushed its tag. If Cloud Run deploy runs too early, it can fail with
Image ... not found; rerun./scripts/deploy-runtime.shafter the image is visible in Artifact Registry. - A successful Cloud Build does not guarantee that Cloud Run can import an
extremely large runtime image within the deployment deadline. In one
deepseek-r1:671battempt, Cloud Run spent about 2 hours importing the image and then failed the revision withDeadline exceeded. - For giant baked images, the next bottleneck after Cloud Build can be Cloud Run image import time rather than container startup time. If a revision never reaches startup logs and the status mentions image import progress, consider a smaller model or a different runtime platform.
- For very large models, use low concurrency.
CONCURRENCY="1"is the safest starting point. MODEL_CONTEXT_WINDOWis a client/profile declaration for Codex, not a hard guarantee that the runtime can serve that full context under memory pressure.
The following profile is the current known-good end-to-end deployment in this repo for a Cloud Run GPU service plus private proxy:
GCP_REGION="europe-west4"
SERVICE_PREFIX="mixtral-8x22b"
OLLAMA_MODEL="mixtral:8x22b"
MODEL_ALIAS="mixtral-8x22b"
GPU_TYPE="nvidia-rtx-pro-6000"
CLOUD_RUN_CPU="30"
CLOUD_RUN_MEMORY="120Gi"
MAX_INSTANCES="1"
CONCURRENCY="1"
STARTUP_INITIAL_DELAY_SECONDS="240"
PROXY_CPU="2"
PROXY_MEMORY="2Gi"
MODEL_CONTEXT_WINDOW="32768"
REQUEST_TIMEOUT_SECONDS="7200"Operational notes for this validated Mixtral profile:
- Use the proxy service as the client entry point. The runtime service is not the public API surface.
- The model works through the OpenAI-compatible proxy endpoints and Codex can
authenticate with
gcloud auth print-identity-token. - With
MAX_INSTANCES="1"andCONCURRENCY="1", back-to-back or parallel requests can hit Cloud Run capacity limits while the single GPU instance is cold-starting or busy. If lower latency matters more than cost, consider settingmin-instances=1on the runtime.
The following profile documents the earlier large-model experiment with
deepseek-r1:671b on nvidia-rtx-pro-6000. It is included as a sizing note,
not as a verified working deployment for this repo:
GCP_REGION="europe-west4"
SERVICE_PREFIX="deepseek-r1"
OLLAMA_MODEL="deepseek-r1:671b"
MODEL_ALIAS="deepseek-r1-671b-iq4xs"
GPU_TYPE="nvidia-rtx-pro-6000"
CLOUD_RUN_CPU="30"
CLOUD_RUN_MEMORY="120Gi"
MAX_INSTANCES="1"
CONCURRENCY="1"
STARTUP_INITIAL_DELAY_SECONDS="240"
MODEL_CONTEXT_WINDOW="4096"
REQUEST_TIMEOUT_SECONDS="7200"For that profile, the runtime image build should also use a larger Cloud Build worker than the original template defaults. A practical starting point is:
options:
machineType: E2_HIGHCPU_32
diskSizeGb: 1000
timeout: 86400s-
Sign in to Google Cloud:
gcloud auth login gcloud auth application-default login
-
Create your config:
cp config.example.env config.env
-
Edit
config.env.Set
GCP_PROJECT_ID, chooseGCP_REGION, and set any Ollama model you want:OLLAMA_MODEL="qwen3-coder:30b" MODEL_ALIAS="qwen3-coder-30b"
You are not required to use
gpt-oss. Any Ollama model name is valid as long as the selected Cloud Run CPU, memory, GPU, and region can run it.If you want the tested RTX PRO 6000 profile from this repo, start from the "Verified Mixtral-8x22b Profile" above. The DeepSeek section is a record of a failed large-model attempt, not the default recommendation.
-
Prepare Google Cloud resources:
./scripts/bootstrap-gcp.sh
-
Build the model image:
./scripts/build-model-image.sh
This step waits for Cloud Build to finish, but very large models can still take hours. For example,
deepseek-r1:671btook about 7 hours 35 minutes in one run. Wait for this command to complete successfully before deploying the runtime. -
Deploy the private Ollama runtime:
./scripts/deploy-runtime.sh
-
Deploy the private proxy:
./scripts/deploy-proxy.sh
-
Add a Codex profile:
./scripts/write-codex-profile.sh --install
-
Smoke test the normal API and Codex:
./scripts/smoke-test.sh
Codex or HTTP client
|
| Cloud Run IAM identity token
v
Private proxy service
|
| service account identity token
v
Private Ollama runtime service
The runtime is not meant to be called directly. The proxy is the stable API surface for both normal clients and Codex.
Fetch an identity token and call the proxy URL:
TOKEN="$(gcloud auth print-identity-token)"
source scripts/lib/config.sh
PROXY_URL="$(gcloud run services describe "$(proxy_service_name)" \
--project "$GCP_PROJECT_ID" \
--region "$GCP_REGION" \
--format 'value(status.url)')"
curl -sS "$PROXY_URL/v1/responses" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"model":"mixtral-8x22b","input":"Respond exactly: OK"}'Chat completions also work:
curl -sS "$PROXY_URL/v1/chat/completions" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{"model":"mixtral-8x22b","messages":[{"role":"user","content":"Respond exactly: OK"}]}'OLLAMA_MODEL is the model pulled during Cloud Build. Examples:
qwen3-coder:30bmixtral:8x22bgpt-oss:120bllama3.1:8bdeepseek-coder:latest
MODEL_ALIAS is what the proxy exposes to clients and Codex. It can match
OLLAMA_MODEL, but a colon-free alias is usually quieter in Codex telemetry.
Large models need larger images, longer Cloud Build time, enough GPU VRAM, and
larger Cloud Run CPU/memory. Tune CLOUD_RUN_CPU, CLOUD_RUN_MEMORY,
GPU_TYPE, MAX_INSTANCES, and CONCURRENCY for your model.
For very large models on Ollama, also tune Cloud Build worker size, build disk, and build timeout. Build-time model download size can be much larger than the final deployment shape suggests.
Services deploy with --no-allow-unauthenticated.
- Users call the proxy with a Google identity token.
- The proxy calls the runtime with its service account identity token.
- The runtime service is not public.
This template does not include public API keys in V1.