Run a 125-billion-parameter AI model on a normal gaming PC
one NVIDIA card (12-24 GB) + 64 GB of RAM · Windows or Linux · one click to install
Strata runs Qwen3.8-Flash-Next - a large, smart AI model that normally needs a server - on your own PC. It writes its answers at 60-95 tokens per second (a token is about ¾ of a word): faster than you can read.
- Private: everything runs on your PC. Nothing is sent anywhere.
- Works with your apps: chat apps, coding agents and scripts that speak the OpenAI or Anthropic API just work.
- Sees pictures too, if you want (screenshots, photos, scanned pages).
- Free and open source.
Jump to: Is my PC enough? · Install · How fast? · Which model? · Using it · Problems? · How it works · All the details
| You need | |
|---|---|
| Graphics card | NVIDIA RTX 30, 40 or 50 series with 12 GB of VRAM or more |
| Memory (RAM) | 64 GB |
| Free disk space | ~80 GB (an SSD makes the first start much faster) |
| System | Windows 10/11, or Linux |
That's it. The only thing you install yourself is a current NVIDIA driver (nvidia.com/drivers or the NVIDIA App). Everything else - Python, the engine, the model - is set up for you.
Windows
- Download this project and unzip it (or
git cloneit). - Double-click
START-HERE.bat. - Answer 4 questions - or just press Enter each time for the recommended choice:
- Which model? The original, or Swift 1.5 (a version that thinks shorter and answers sooner)
- Which size? Q2_0, IQ2_XS, IQ3_XXS or IQ3_S - see which model
- How much context? How much text it can keep in mind at once (it suggests one for your card)
- Images? Whether it should also read pictures
Then it downloads everything (the model is ~70 GB, so the first time takes a while - you can stop and it picks up
where it left off) and starts the model. Your browser opens the Strata app at http://127.0.0.1:8080.
Next time, just double-click START-HERE.bat again: it starts right away, nothing is downloaded twice. Close its
window to stop the model.
Linux: run ./setup.sh - same questions, same result.
Measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM:
| Size | Writes answers (short chat) | Writes answers (128K context) | Reads your prompt |
|---|---|---|---|
| Q2_0 | 95 tokens/s | 65 tokens/s | 539 tokens/s |
| IQ2_XS | 78 tokens/s | 52 tokens/s | 463 tokens/s |
| IQ3_XXS | 66 tokens/s | 45 tokens/s | 410 tokens/s |
| IQ3_S | 54 tokens/s | 42 tokens/s | 374 tokens/s |
- Writes answers = how fast the reply appears (tokens per second).
- Reads your prompt = how fast it takes in what you send (long documents, code, chat history).
A card with more VRAM is faster, because more of the model fits on the GPU: an RTX 3090 (24 GB) should do roughly 100-140 tokens per second. All measurements, long-context numbers and estimates for other cards are in the details.
The size (the same model, compressed more or less):
| Size | Download | Speed | Quality | Pick it if... |
|---|---|---|---|---|
| Q2_0 | 66 GB | fastest | good | you want speed |
| IQ2_XS | 68 GB | fast | better | you want a good all-rounder (recommended) |
| IQ3_XXS | 76 GB | slower | great | you want better answers (uses 43 GB of your 64 GB RAM) |
| IQ3_S | 84 GB | slowest | best: matches the full model on the published tests | you want the very best answers (original model only; uses 50 GB of your 64 GB RAM, so close other big programs) |
The version:
- Qwen3.8-Flash-Next - the original.
- Swift 1.5 - a fine-tune by UkisAI that thinks much shorter before answering, so you get the answer sooner, with about the same quality. Same speed per token. Its own license applies (see its page).
Not sure? Take IQ2_XS. You can add another one later with START-HERE.bat --setup.
- In the browser:
http://127.0.0.1:8080- the Strata app (it opens by itself when the model starts): Chat, a live Monitor of the model and your GPU/CPU/RAM, and About with the settings and addresses. - Chat in the terminal:
.venv\Scripts\python chat.py - Your apps and coding agents: add it as an "OpenAI-compatible" provider with base URL
http://127.0.0.1:8080/v1, any API key and any model name. Apps that use Anthropic's API:http://127.0.0.1:8080/v1/messages. - Thinking: the model thinks before it answers. Choose off, low, medium or high - in the chat page menu, with
/think lowinchat.py, or with your app's "reasoning effort" setting. Off is fastest; high is best for hard questions. - Pictures: in the chat page click Picture; in
chat.pytype/image <path>; in apps just attach them. - From your phone or another PC: see the details (set an API key first).
Good to know: it answers one request at a time. The first message of a chat is read in full (about 1 minute per 30,000 tokens); after that it keeps the conversation and reads only what is new, so follow-ups start in seconds.
| What you see | What to do |
|---|---|
the NVIDIA driver is too old |
Update the driver (NVIDIA App or nvidia.com/drivers), restart the PC, run START-HERE.bat again. |
| It stopped during download or setup | Run START-HERE.bat again - it continues where it stopped. |
port 8080 is already in use |
Strata is already running - look for its window. |
| The first start takes minutes | Normal: it loads 35-43 GB into RAM. The next start is faster. |
| Slow, and the disk light is busy | Not enough free RAM: close other programs (browsers use a lot), or pick Q2_0 / IQ2_XS. |
| "prompt exceeds the context" | The conversation is longer than the context you chose: run START-HERE.bat --setup and pick more. |
More in the full troubleshooting table. Still stuck? Open an issue and attach
strata-<model>.log from this folder.
A model this big doesn't fit on a gaming graphics card. Strata splits the work between the parts of your PC:
- The GPU runs the part of the model that is used for every word, plus the "experts" it needs most often.
- The RAM holds all 24,576 experts, and the CPU computes the few the GPU doesn't have - at the same time as the GPU.
- The SSD holds a big lookup table; the model reads a few rows of it per word.
- A small helper inside the model guesses the next words, and Strata checks several guesses at once. That makes it 1.6-1.8x faster than going word by word - and the answer is exactly the same.
The full story is in the paper and the details.
- Model: Qwen3.8-Flash-Next by the Qwen team; compressed versions by ISTA-DASLab; Swift 1.5 by UkisAI. Their licenses apply to the model files.
- Built with parts of llama.cpp / ggml (MIT). Ideas from Splash, ninfer and HyperQwen. More in the details.