Skip to content

feat: independently grade robot evidence and expose all GPU agent trials - #10

Merged
noteflowai merged 3 commits into
mainfrom
codex/research-rollout-20260914
Sep 14, 2026
Merged

noteflowai merged 3 commits into
mainfrom
codex/research-rollout-20260914

Conversation

@noteflowai

Copy link
Copy Markdown
Owner

Agent workflow completion and skill delivery do not establish task acceptance. Add a robot-evidence-review task with independently derived numerical checks and six fault controls, a bounded non-root model workspace, separate Docker startup readiness, and Harbor/ATIF interoperability.

The new homepage explorers retain 27 real skill-delivery trials, 12 composition trials and six cross-model handoffs, including all failures, candidate files, original receipts and independent grades. Source-checked manifests reject invented headlines even when rehashed. Funes retrieval and native Harbor oracle/NOP execution are recorded; no efficacy gain or full-benchmark claim is made.

Validation: 302 tests; Ruff check/format; independent installed-wheel checks for both robot task languages; real Python/JavaScript Docker fault audits; 33 exports validated by Harbor0.23.0's ATIF schema; browser checks at1440/390/320 including recovery and shared state. Version0.9.0. Public immutable dataset: https://huggingface.co/datasets/glayguo/noteflow-research-pilots/tree/v2026-09-14.

@noteflowai
noteflowai marked this pull request as ready for review September 14, 2026 12:25
@noteflowai
noteflowai merged commit f10cd61 into main Sep 14, 2026
12 checks passed
@noteflowai
noteflowai deleted the codex/research-rollout-20260914 branch September 14, 2026 12:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant