JARVIS coordinates a handful of free and very cheap LLMs — NVIDIA NIM, OpenRouter, DeepInfra, Groq, Gemini — to approximate frontier-model coding quality for a few cents a task, just slower. The bet behind it: instead of one weak model thinking alone, have several models plan independently, critique each other, merge the best plan, implement it, and then verify the result by running it — and push as much of the hard, error-prone bookkeeping as possible into a deterministic harness so the weak model only ever has to make a small, local decision.
The most developed part — and the part worth using — is the coding agent ("Deep Code"): a four-stage pipeline (UNDERSTAND → PLAN → IMPLEMENT → REVIEW) aimed at solving real GitHub issues on real repositories. On the SWE-bench Pro Python set it resolves ~63% of a hard tuned subset and ~52% of a held-out set it was never tuned against, and on the SWE-bench Pro Go set it now resolves ~55% of a 20-instance subset — from models that individually are nowhere near that level (see Benchmarks).
This is a research project, not a finished product. Be selective:
- ✅ Use the coding agent (Deep Code). This is where almost all the engineering went. It plans, edits real files, runs its own checks, and ships a patch. It's the real deal.
- 🧪
medium_chatis a promising experiment — a new multi-brain chat backbone (3 proposers + a merger). It works surprisingly well and is opt-in; see The experimental chat backbone. ⚠️ Everything else works but is far from optimal — the legacy general chat, web search, image gen, formatting. They run, but they haven't had the attention the coder has.- 🚫 Don't rely on the auto-router. The classifier that decides "is this chat / a question / a coding task" is the weakest link. Skip it: go straight to Deep Code.
The recommended way to use JARVIS:
python ui_main.py # opens the web UIThen in the UI: paste your free API keys in Settings → API Keys, set your project path in Settings → Project, and run your coding task in Deep Code mode. That's the path that's been tuned and tested.
API keys are read from env vars (or the UI settings): OPENROUTER_API_KEY(S), NVIDIA_API_KEY, DEEPINFRA_API_KEY, GROQ_API_KEY, GEMINI_API_KEY(S). You need OpenRouter at minimum; the rest add redundancy.
A single free/weak model is decent but not frontier-level, and — more importantly — it's unreliable: it loses track of indentation, edits the wrong line, forgets which file it's in, reinvents a format. JARVIS's whole design is two ideas stacked together.
UNDERSTAND → PLAN → IMPLEMENT → REVIEW
(2 drafts (one coder, (run a repro,
→ 1 merger) JSON-ops) route a fix)
- PLAN — two planners draft a plan independently and in parallel (deepseek-v4-flash and gpt-oss-120b — deliberately diverse, open, non-frontier models). A merger (deepseek-v4-flash) then consolidates them into one plan, taking the most correct and complete approach rather than the longest. Disagreement between drafts is a feature: it surfaces the parts that are actually hard.
- IMPLEMENT — the coder (deepseek-v4-flash) executes one plan step at a time through structured JSON-ops tool calls —
read_file,edit_file,create_file,search_text,find_refs,run_code,finish, etc. — not by dumping a blob of text. - REVIEW (self-verify) — after the patch, JARVIS writes precise behavioural checks (and/or a small reproduction) from the task's own description, runs the edited code, and if a required behaviour fails — or the patch broke one that worked before — routes a corrective fix back to the coder. A check is only trusted when it flips fail→pass between the original and edited tree, so brittle checks can't cause a false approval. It never reads the project's hidden test suite — that would be cheating; it tests what the task actually says. Works on any repo — a homemade project (local bwrap sandbox + your active venv's deps) as well as a benchmarked one (the instance's Docker image). Opt-in: enable with
python3 main.py --verifyorJARVIS_ENABLE_REVIEW=1.
The whole pipeline runs on deepseek-v4-flash (pinned to the official DeepSeek provider first on OpenRouter for speed), with gpt-oss-120b as a second, diverse planner voice.
This is the idea that makes weak models reliable. Anything global, stateful, or easy to get wrong is handled deterministically by the harness, so the model is left with a single small move:
- Edits are number-first and content-verified. The coder copies a line from the view (which carries both its line number and its content) and the harness applies the change by anchoring on both — so a stale line number self-corrects and a wrong anchor is rejected instead of silently corrupting the file.
- Indentation is computed, not guessed. The file view encodes each line's indent as a number; the harness re-emits the real whitespace. The model never counts spaces.
- Edits run through safety gates before they land: a parse check (reject + revert on a syntax error the edit introduced); an undefined-name / dangling-reference check (you can't delete a symbol that's still called); a final-state dangling-call gate (if the finished patch calls a symbol that's defined nowhere — because a later edit silently dropped its definition — it routes back to re-add it); a plan step-integrity guard (a distinct plan step is never lost to a duplicate number); an interface-coverage oracle (every function the issue names as a target must actually be modified); and an apply-time guard that refuses to overwrite a working file with empty content.
- The code map, symbol lookups, stale-read detection, and "which file am I in" are all the harness's job. The model asks; the harness answers with ground truth.
The slogan, from the project's own notes: the harness computes the global, the model acts local.
Much of this repo's history is hard-won, offline-verified fixes to make weak models behave. The big themes:
- A reflex library for the coder. Instead of vague advice, the coder's prompt carries concrete, triggered reflexes drawn from real observed failures: read from the source you gated on; "all / every / collect" means accumulate, don't overwrite; produce the exact type/literal a test expects; bytes stay bytes until you decode them; remove a symbol and fix every call site in one edit; a missing third-party import is the environment, not your bug. (A repeatedly-validated lesson: for weak models, concrete reflexes beat elegant abstract principles — when we tried replacing them with general principles, the score dropped.)
- Delivery-loss guards. A whole class of failures turned out to be the machinery silently discarding the model's correct work — a dropped plan step, a leaked tool-call blob that collapsed a plan, a self-verify revert that discarded a needed definition, a defined method clobbered by a later edit. These are now caught deterministically. On a held-out audit, every remaining failure was a genuine reasoning/contract miss — zero were machinery bugs.
- A self-verifying reviewer under a strict snapshot-and-revert invariant: the review can only help or be neutral, never ship a patch worse than the coder's original. The repro author sees the changed symbols' signatures only — not the implementation — so it can't be primed into rubber-stamping a buggy patch.
- Robust, fast provider routing. deepseek-v4-flash is pinned to the official DeepSeek endpoint (fastest, and its prompt cache makes the growing file-view nearly free); API keys are round-robined, dead/over-quota keys are skipped, and 402/429 responses are retried down the chain. A single provider hiccup doesn't sink a run.
- A read-only, no-network sandbox (bwrap) where all edits land and
run_codeexecutes, so nothing the agent does touches your real files until you approve.
A full 21-instance held-out benchmark run — plan + implement + review, several model calls per instance — cost $1.25 total (~6¢ an instance). The whole pipeline runs on deepseek-v4-flash, whose provider-side prompt cache makes the biggest expense (the ever-growing file view the coder reads each round) nearly free after the first round. So you're paying cents for a full multi-stage run, not dollars.
(Earlier versions ran fully $0 on free tiers; the current setup spends a little to get a much more reliable, much faster coder.)
JARVIS is developed against SWE-bench Pro — real GitHub issues, graded by actually running the project's hidden tests in Docker (the ScaleAI harness + per-instance images). Pro is multi-language (Python, Go, TypeScript/JavaScript) and ships the issue's behavioral spec, which is a fair, hard target. Every number below is graded by real Docker pass/fail.
| Set | Score | What it is |
|---|---|---|
| Python — tuned subset (27 instances) | 17/27 ≈ 63% | the hard set we developed against, reproducible across runs |
| Python — held-out (21 instances) | 11/21 ≈ 52% | brand-new instances the fixes were never tuned against |
| Go — tuned subset (20 instances) | 11/20 ≈ 55% | the set the Go work was tuned against; the jump came from a general merger fix (see below), not Go hacks — a held-out Go check is in progress |
For the Python set, the held-out number is the one that matters, and it's the honest one. Getting ~52% on instances we'd never seen — versus ~63% on the set we tuned — is a small, expected gap, and it's well above the "we overfit" threshold. In other words: the improvements are real capability, not memorization of the tuning set. We verified this deliberately, on a separate subset, precisely so the headline number can't be a cherry-pick.
Two more honest points:
- A follow-up held-out run (with the newest delivery-loss guards) was in progress and trending above 52% when we stopped it — it was already good enough to make the point, and improvements were still landing. So ~52% is a floor for the held-out generalization, not a ceiling.
- A forensic on the held-out failures found that of 10 misses, only 2 were genuinely impossible (the hidden test pins an exact string/signature not derivable from the issue text); the rest were avoidable reasoning errors. The realistic reachable ceiling on this set is ~76%, and — notably — every failure was the model producing a complete-but-wrong patch, not the harness losing correct work. The machinery is solved; what's left is semantic quality.
Separately, as a real-world sanity check we had JARVIS build and then iteratively extend a small application from scratch over many rounds (add features, fix bugs, refactor), running the app after each round. It built a working CLI app, debugged a regression from just a symptom report, and shipped features end-to-end — surfacing bugs that pure benchmark runs never would.
The same pipeline runs on Go, and on a 20-instance SWE-bench Pro Go subset it now resolves 11/20 ≈ 55% (on the low-variance instances the stable line moved from a rock-solid 5/13 to 7/13). The jump did not come from Go-specific hacks — it came from fixing a general reasoning bug in the plan-merger.
The merger's job is to take the independent draft plans, debate them, and merge the best one. It was rubber-stamping: it would accept a draft's claim that something was "already handled" — an output field already produced, a caller already compatible, a code path already covered — without ever reading the code to check. On real issues that assumption is often wrong (the caller still uses the old return type; the field is produced for the old shape, not the new one), so the merged plan silently dropped a required step and the patch failed the hidden test.
The fix was to make the merger verify instead of assume: any "already handled / already works / unchanged" claim must be confirmed by reading the code that supposedly handles it — and if it isn't read, it becomes a real plan step, never a silent tick. Combined with trimming the merger prompt down so its decisive checks actually bind (a bloated instruction wall is one a weak model reads and drops), this made the merger stop rushing to a confident wrong answer.
Two things followed, and both are the point:
- It scores higher (the stable line moved +2), and we watched the mechanism — on a teleport instance it had previously failed on, the merger now read and included the client caller (
weblogin.go) it used to wave away, and the instance passed. - It scores more reliably — less variance, because "commit to an idea and check it" replaces "assume it's fine and hope."
Because the merger is language-agnostic, this shouldn't be a Go-only win: it's the same merger the Python pipeline uses, so we expect the fix to help there too. (Honest caveat: the Python numbers above predate this fix and haven't been re-measured yet, and a held-out Go run — different instances, never tuned against — is in progress to confirm the Go result isn't overfit to the tuning subset.)
Multi-language is the next phase. The pipeline above was first tuned for Python; Go is now at ~55% on its subset, and Pro also includes TypeScript/JavaScript. JARVIS already has per-language detection + per-step reflex injection. Extending the safety gates and prompts to full parity across 20+ languages is in progress.
medium_chat applies the same multi-brain idea to conversation — the thing you talk to, rather than the coder. It's opt-in and experimental, but it works well (it holds up strongly against a frontier-model judge on hard prompts).
Each turn runs a propose → merge → execute cycle:
- PROPOSE (parallel): three proposers — gemma, gpt-oss-120b, and deepseek-v4-flash — each look at the retrieved context + conversation + prior tool results and independently propose a step (
{reasoning, tools:[…]}). Free tiers first, escalating to paid only on error. - MERGE: a merger (deepseek-v4-flash) consolidates the three proposals into one final step, preferring tools that ≥2 proposers agreed on (with a fallback to the best single proposal if it errors). Independent agreement is the signal.
- EXECUTE: the chosen tools run. The only user-visible output is a
message_usertool — the model reasons in free prose and emits actions as plain markers. Tools includeweb_search,fetch(it can read a URL or an arXiv PDF — AI-driven browsing, not just snippets), andrun_command(through thecore/cmd_execsandbox: safe commands auto-run, dangerous ones are refused pending confirmation, blocked ones never run). The turn ends when amessage_useris emitted and nothing needs a follow-up; otherwise tool results feed back and the loop continues.
Deterministic gates keep the weak models honest (e.g. a no-action reply that promises to search/fetch/run is bounced back to actually do it; raw-text answers with LaTeX/quotes/matrices are handled without mangling). Context comes from the context_rag retrieval extender.
How to try it:
JARVIS_MEDIUM_CHAT=1 python ui_main.py # whole session
# or, per-message in either the CLI or the web UI, prefix:
!!medium_experimental your message heremain.py— terminal CLI entry pointui_main.py— web UI (recommended)workflows/code.py— the coding pipeline (plan → implement → review)workflows/medium_chat.py— the experimental multi-brain chat backbonecore/native_tools.py— the coder's tool loop (JSON-ops + native function calling)core/self_verify.py— the self-verifying reviewercore/prompts_v8.py— the live prompts (planner / coder / reviewer)core/lang_detect.py— deterministic per-file/per-codebase language detectioncontext_rag/— retrieval extender used by the chat backbonetools/— code index, sandbox, codebase viewsclients/— provider routing and fallback
JARVIS is an experiment in getting frontier-ish coding out of models that, alone, aren't close — by making them collaborate and by doing the hard bookkeeping for them. The coding agent is the part that delivers on that today; the rest is catching up.