Skip to content
 
 

Repository files navigation

Strata

Run a 125-billion-parameter AI model on a normal gaming PC
one NVIDIA card (12-24 GB) + 64 GB of RAM · Windows or Linux · one click to install

Strata runs Qwen3.8-Flash-Next - a large, smart AI model that normally needs a server - on your own PC. It writes its answers at 60-95 tokens per second (a token is about ¾ of a word): faster than you can read.

  • Private: everything runs on your PC. Nothing is sent anywhere.
  • Works with your apps: chat apps, coding agents and scripts that speak the OpenAI or Anthropic API just work.
  • Sees pictures too, if you want (screenshots, photos, scanned pages).
  • Free and open source.

Jump to: Is my PC enough? · Install · How fast? · Which model? · Using it · Problems? · How it works · All the details


Is my PC enough?

You need
Graphics card NVIDIA RTX 30, 40 or 50 series with 12 GB of VRAM or more
Memory (RAM) 64 GB
Free disk space ~80 GB (an SSD makes the first start much faster)
System Windows 10/11, or Linux

That's it. The only thing you install yourself is a current NVIDIA driver (nvidia.com/drivers or the NVIDIA App). Everything else - Python, the engine, the model - is set up for you.

Install (3 steps)

Windows

  1. Download this project and unzip it (or git clone it).
  2. Double-click START-HERE.bat.
  3. Answer 4 questions - or just press Enter each time for the recommended choice:
    • Which model? The original, or Swift 1.5 (a version that thinks shorter and answers sooner)
    • Which size? Q2_0, IQ2_XS, IQ3_XXS or IQ3_S - see which model
    • How much context? How much text it can keep in mind at once (it suggests one for your card)
    • Images? Whether it should also read pictures

Then it downloads everything (the model is ~70 GB, so the first time takes a while - you can stop and it picks up where it left off) and starts the model. Your browser opens the Strata app at http://127.0.0.1:8080.

Next time, just double-click START-HERE.bat again: it starts right away, nothing is downloaded twice. Close its window to stop the model.

Linux: run ./setup.sh - same questions, same result.

How fast is it?

Measured on an RTX 5070 (12 GB), a Ryzen 5 7600 and 64 GB of RAM:

Size Writes answers (short chat) Writes answers (128K context) Reads your prompt
Q2_0 95 tokens/s 65 tokens/s 539 tokens/s
IQ2_XS 78 tokens/s 52 tokens/s 463 tokens/s
IQ3_XXS 66 tokens/s 45 tokens/s 410 tokens/s
IQ3_S 54 tokens/s 42 tokens/s 374 tokens/s
  • Writes answers = how fast the reply appears (tokens per second).
  • Reads your prompt = how fast it takes in what you send (long documents, code, chat history).

A card with more VRAM is faster, because more of the model fits on the GPU: an RTX 3090 (24 GB) should do roughly 100-140 tokens per second. All measurements, long-context numbers and estimates for other cards are in the details.

Which model should I pick?

The size (the same model, compressed more or less):

Size Download Speed Quality Pick it if...
Q2_0 66 GB fastest good you want speed
IQ2_XS 68 GB fast better you want a good all-rounder (recommended)
IQ3_XXS 76 GB slower great you want better answers (uses 43 GB of your 64 GB RAM)
IQ3_S 84 GB slowest best: matches the full model on the published tests you want the very best answers (original model only; uses 50 GB of your 64 GB RAM, so close other big programs)

The version:

  • Qwen3.8-Flash-Next - the original.
  • Swift 1.5 - a fine-tune by UkisAI that thinks much shorter before answering, so you get the answer sooner, with about the same quality. Same speed per token. Its own license applies (see its page).

Not sure? Take IQ2_XS. You can add another one later with START-HERE.bat --setup.

Using it

  • In the browser: http://127.0.0.1:8080 - the Strata app (it opens by itself when the model starts): Chat, a live Monitor of the model and your GPU/CPU/RAM, and About with the settings and addresses.
  • Chat in the terminal: .venv\Scripts\python chat.py
  • Your apps and coding agents: add it as an "OpenAI-compatible" provider with base URL http://127.0.0.1:8080/v1, any API key and any model name. Apps that use Anthropic's API: http://127.0.0.1:8080/v1/messages.
  • Thinking: the model thinks before it answers. Choose off, low, medium or high - in the chat page menu, with /think low in chat.py, or with your app's "reasoning effort" setting. Off is fastest; high is best for hard questions.
  • Pictures: in the chat page click Picture; in chat.py type /image <path>; in apps just attach them.
  • From your phone or another PC: see the details (set an API key first).

Good to know: it answers one request at a time. The first message of a chat is read in full (about 1 minute per 30,000 tokens); after that it keeps the conversation and reads only what is new, so follow-ups start in seconds.

Something went wrong?

What you see What to do
the NVIDIA driver is too old Update the driver (NVIDIA App or nvidia.com/drivers), restart the PC, run START-HERE.bat again.
It stopped during download or setup Run START-HERE.bat again - it continues where it stopped.
port 8080 is already in use Strata is already running - look for its window.
The first start takes minutes Normal: it loads 35-43 GB into RAM. The next start is faster.
Slow, and the disk light is busy Not enough free RAM: close other programs (browsers use a lot), or pick Q2_0 / IQ2_XS.
"prompt exceeds the context" The conversation is longer than the context you chose: run START-HERE.bat --setup and pick more.

More in the full troubleshooting table. Still stuck? Open an issue and attach strata-<model>.log from this folder.

How does it work?

A model this big doesn't fit on a gaming graphics card. Strata splits the work between the parts of your PC:

how Strata splits the model between GPU, RAM and SSD

  • The GPU runs the part of the model that is used for every word, plus the "experts" it needs most often.
  • The RAM holds all 24,576 experts, and the CPU computes the few the GPU doesn't have - at the same time as the GPU.
  • The SSD holds a big lookup table; the model reads a few rows of it per word.
  • A small helper inside the model guesses the next words, and Strata checks several guesses at once. That makes it 1.6-1.8x faster than going word by word - and the answer is exactly the same.

The full story is in the paper and the details.

Credits

About

Qwen3.8-Flash-Next (125B MoE) on a 12-24 GB NVIDIA GPU + 64 GB RAM: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages