IT Architect — corporate banking architecture · independent LLM/RL researcher-practitioner ИТ-архитектор (корпоративная архитектура банка) · независимый исследователь-практик LLM/RL
A domain-tuned agent harness beats a universal one on its own tasks. Доменно-заточенный агентный харнесс бьёт универсальный — на задачах самого универсала.
Everything below is a working proof — built, measured and dogfooded daily. Всё ниже — рабочие доказательства: построено, измерено и используется каждый день.
| Repo | What it is · Что это |
|---|---|
| theseus | Industrial-grade agentic TUI harness for the ML/RL domain — built after a code review of three industry leaders, then deliberately bent where leaders don't go. ~57K LOC · 74 modules · 31 agent tools · 1,367 green tests · 981 skills (1,093 SKILL.md) · 0 clippy warnings. Runs a local GRPO-tuned Qwen3.5-4B as its «Ariadne's thread». |
| spine | Domain harness for solution architects: spine-invariants, ADR workflow, LLM-judge rubrics, skills/plugins, subagents, governance. |
| spine-bank | Spine Banking Edition (arch-be) — harness for banking solution architects; public MIT snapshot of the core. |
| spine-aiml | Spine AI/ML Edition (arch-ml) — meta-harness for AI/ML researchers riding on top of coding harnesses: invariants, fitness gates, handoff contracts. |
| spine-be-distrib | arch-be distribution — binary releases and install guides. |
| Repo | What it is · Что это |
|---|---|
| ariadna-training | Qwen3.5-4B CPT + SFT + GRPO pipeline for ML-concept curation, fed by a private 18.6K-paper arXiv library distilled into concept cards. |
| qwen35-4b-ariadna-grpo | GRPO post-training: 5 iterations (v6–v10), 3-level reward + GRM judge. Weights: GGUF on Hugging Face. |
| alfworld-grpo-siri-reskill | Multi-turn GRPO on ALFWorld — SIRI+Replay skill internalization (Qwen2.5-1.5B). |
| Repo | What it is · Что это |
|---|---|
| agent-eval-skills | Agent Skills for evaluating LLM agents and LLM apps: classification, rubrics, gates. |
| tui-trace-control | Controlling TUI agents' reasoning and actions via local traces — a full 8-skill lifecycle (setup → watch → audit → contain → resume). |
| ai-detector-evasion-skills | How AI-content detectors work, where they fail, and the ethics boundaries — Agent Skills. |
| ai-native-sdlc-architect-artifacts | AI-native SDLC: architectural artifacts for corporate & solution architects. |
| spine-contour-bench | Benchmark of the contour around an agent: gates, traceability, handoff packages. |
| platformv-arch-bench | Benchmark of architecture tasks: harness role & Spine format vs the universal approach. |


