Skip to content
View romannekrasovaillm's full-sized avatar

Block or report romannekrasovaillm

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
romannekrasovaillm/README.md
Typing SVG

Roman Nekrasov

IT Architect — corporate banking architecture · independent LLM/RL researcher-practitioner ИТ-архитектор (корпоративная архитектура банка) · независимый исследователь-практик LLM/RL

Profile views Followers Stars


🧭 The thesis · Тезис

A domain-tuned agent harness beats a universal one on its own tasks. Доменно-заточенный агентный харнесс бьёт универсальный — на задачах самого универсала.

Everything below is a working proof — built, measured and dogfooded daily. Всё ниже — рабочие доказательства: построено, измерено и используется каждый день.

🏛️ Agent harnesses (Rust)

Repo What it is · Что это
theseus Industrial-grade agentic TUI harness for the ML/RL domain — built after a code review of three industry leaders, then deliberately bent where leaders don't go. ~57K LOC · 74 modules · 31 agent tools · 1,367 green tests · 981 skills (1,093 SKILL.md) · 0 clippy warnings. Runs a local GRPO-tuned Qwen3.5-4B as its «Ariadne's thread».
spine Domain harness for solution architects: spine-invariants, ADR workflow, LLM-judge rubrics, skills/plugins, subagents, governance.
spine-bank Spine Banking Edition (arch-be) — harness for banking solution architects; public MIT snapshot of the core.
spine-aiml Spine AI/ML Edition (arch-ml) — meta-harness for AI/ML researchers riding on top of coding harnesses: invariants, fitness gates, handoff contracts.
spine-be-distrib arch-be distribution — binary releases and install guides.

🧪 Post-training line (CPT → SFT → GRPO / RLVR)

Repo What it is · Что это
ariadna-training Qwen3.5-4B CPT + SFT + GRPO pipeline for ML-concept curation, fed by a private 18.6K-paper arXiv library distilled into concept cards.
qwen35-4b-ariadna-grpo GRPO post-training: 5 iterations (v6–v10), 3-level reward + GRM judge. Weights: GGUF on Hugging Face.
alfworld-grpo-siri-reskill Multi-turn GRPO on ALFWorld — SIRI+Replay skill internalization (Qwen2.5-1.5B).

🛠️ Skills, evals & benchmarks

Repo What it is · Что это
agent-eval-skills Agent Skills for evaluating LLM agents and LLM apps: classification, rubrics, gates.
tui-trace-control Controlling TUI agents' reasoning and actions via local traces — a full 8-skill lifecycle (setup → watch → audit → contain → resume).
ai-detector-evasion-skills How AI-content detectors work, where they fail, and the ethics boundaries — Agent Skills.
ai-native-sdlc-architect-artifacts AI-native SDLC: architectural artifacts for corporate & solution architects.
spine-contour-bench Benchmark of the contour around an agent: gates, traceability, handoff packages.
platformv-arch-bench Benchmark of architecture tasks: harness role & Spine format vs the universal approach.

📊 GitHub stats

GitHub stats GitHub streak Top languages Activity graph

🧰 Stack

Stack icons

🐍

Contribution snake

Research projects — not for production · Исследовательские проекты, не для прода

Pinned Loading

  1. ariadna-training ariadna-training Public

    Ariadna: Qwen3.5-4B CPT + SFT + GRPO training pipeline for ML concept curation

    Python 1

  2. spine spine Public

    Spine — доменный харнесс solution-архитектора: spine-инварианты, ADR, рубрики с LLM-судьёй, скиллы/плагины, субагенты, губернанс · A domain agent harness for solution architects (Rust TUI/CLI). Исс…

    Rust 1

  3. theseus theseus Public

    Агентный TUI-харнесс на Rust под ML/RL-домен: 28 инструментов, 363 скилла, локальная Ариадна Qwen3.5-4B, 1 321 тест

    Rust 1

  4. agent-eval-skills agent-eval-skills Public

    Набор скиллов для оценки LLM-агентов и LLM-приложений: классификация, рубрики, согласие оценщиков, отчёты

    Python

  5. qwen35-4b-ariadna-grpo qwen35-4b-ariadna-grpo Public

    GRPO post-training for Qwen3.5-4B ML concept explanation. 5 iterations (v6-v10), KAT-Coder-V2.5 inspired 3-level reward + GRM judge.

    Python

  6. spine-bank spine-bank Public

    Spine — архитектурный контур внутри вашего CLI-агента (GigaCode, Claude Code, Kimi, Qwen, omp, OpenClaw): MCP-сервер, 62 скилла (вкл. плейбуки spine-workflows), хуки-гейты, судья без API-ключей. Ил…

    Rust 1