Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
84 commits
Select commit Hold shift + click to select a range
987bccf
feat: clarity integration.
Jul 13, 2026
a2379d6
feat: drive Clarity integration through MCP server instead of CLI.
Jul 14, 2026
42bbfcb
feat: ASSERT + ACS integration.
Jul 17, 2026
e9a8f68
feat: make govern-remeasure loop deterministic and self-explaining.
Jul 18, 2026
f4256f3
test: billing support agent.
Jul 19, 2026
8332b61
fix: increase sample_size to 25 to reduce noise in A/B testing
Jul 19, 2026
2bc166b
fix: delete old eval_config.yaml billing agent.
Jul 19, 2026
4463519
fix: updated agent_guarded.py to source callerid from session.
Jul 19, 2026
ee142d0
fix: give user choice for sample size, add numeric/threshold gate for…
Jul 20, 2026
ac12639
feat: new subsection for output/input points.
Jul 20, 2026
4359195
fix: reduce govern to pure enforcement.
Jul 20, 2026
91460f1
fix: refine SYSTEM_PROMPT to alight with examples.
Jul 20, 2026
49ea571
feat: billing_support_agent resources. clarity, example, and artifacts.
Jul 21, 2026
957da32
fix: forward score_keys in buildJudgedSampleRow to fix false judgefai…
Jul 21, 2026
3e1cdfd
docs: add per-domain organization guidance for multi-run workflows.
Jul 21, 2026
0380f5b
fix: forward score_keys in normalized result items to stop false judg…
Jul 21, 2026
1338d9f
feat(test): azure_doc_qa, change_congrol_agent, travel_planner_langgr…
Jul 23, 2026
b699dff
fix: make fabricated-details output gate history-aware. adopt regen o…
Jul 23, 2026
af69b61
feat: career_health_assessment agent example run.
Jul 23, 2026
fd95607
fix: refine career_health_assessment annotator.
Jul 24, 2026
ade0e69
feat(examples): add Prompt Agents governance packages from SKILL work…
Jul 24, 2026
7d16f3e
fix: update max_turns default for SKILL to 10.
Jul 24, 2026
ddd6dfc
feat(example): science_research_agent and travel_planner_neurosan exa…
Jul 28, 2026
c166795
feat(viewer): headline policy violations split by behavior permissibi…
Aug 1, 2026
a80883e
fix(tests): copy permissibility.ts into the node test harnesses.
Aug 1, 2026
6817a31
fix(examples): delete all skill generated content for rerun of finali…
Aug 1, 2026
5e9a7f0
Merge branch 'main' into ango10/assertacs-skill-bugbash
Aug 1, 2026
c3d7e16
Merge branch 'ango10/viewer-headline-metrics' into ango10/assertacs-s…
Aug 1, 2026
b28812a
feat(example): updated skill and billing_support_agent example run.
Aug 2, 2026
b906117
career_health_assessment: Clarity->ASSERT->ACS->ASSERT bug-bash run
Aug 2, 2026
8d9d990
change_control_agent: Clarity protocol, ACS policy, and governed agent
Aug 2, 2026
0b6a836
prompt_agents: runnable agent, Clarity protocol, ACS policy, governed…
Aug 2, 2026
4220f0a
Add ACS governance for travel_planner_langgraph fabricated travel det…
Aug 2, 2026
299919d
Add ACS governance for azure_doc_qa internal document disclosure
Aug 2, 2026
96b5438
Add ACS governance for science_research_agent restricted content leakage
Aug 2, 2026
1770080
Add ACS governance for travel_planner_neurosan fabricated travel details
Aug 2, 2026
6f1fb0a
Add failure-2 governance for change_control and prompt_agents
Aug 3, 2026
35b5131
Add failure-2 governance for science_research_agent (cross-document i…
Aug 3, 2026
4845b60
Add failure-2 governance for career_health_assessment (CV untrusted-i…
Aug 3, 2026
3553af7
Add failure-2 governance for azure_doc_qa (self-asserted identity esc…
Aug 3, 2026
1f8f222
Add failure-2 governance for travel_planner_neurosan (false budget co…
Aug 3, 2026
5ef87e9
Add failure-2 governance for travel_planner_langgraph (unmarked claim…
Aug 3, 2026
02dc6c2
Update science_research_agent tool cache from phase-1 governed re-run
Aug 3, 2026
b97b503
prompt_agents: add model_only/simtools/gentools variant matrix
Aug 3, 2026
943c867
Reset billing_support_agent to pre-skill state for clean rerun.
Aug 3, 2026
18051f8
feat(agents): update systemize, judge to gpt-5.4, default model to gp…
Aug 3, 2026
cedcdc7
feat(example): billing_support_agent ran through workflow.
Aug 3, 2026
bf95fed
feat(viewer): retire policy_violation/overrefusal from display surfaces.
Aug 3, 2026
d9a5650
fix(cli): detect the permissibility split by key presence, not rate.
Aug 3, 2026
5fc708e
Merge ango10/viewer-headline-metrics into ango10/assertacs-skill-bugbash
Aug 3, 2026
4dccfa3
Merge origin/main into ango10/assertacs-skill-bugbash
Aug 3, 2026
56a2877
feat(example): return billing_support_agent to its pre-skill state.
Aug 3, 2026
9a03901
fix(skill): update SKILL to prevent custom judge dimension generation.
Aug 4, 2026
7d74d61
feat(example): billing_support_agent final workflow demo.
Aug 4, 2026
b0acb99
feat(example): billing_support_agent Clarity Protocol directory moved…
Aug 4, 2026
10505ee
feat(example): career_health_assessment cleared to pre-skill state.
Aug 4, 2026
06edeb9
feat(example): career_health_assessment ran through SKILL workflow.
Aug 4, 2026
0fd51bf
feat(example): cleared azure_doc_qa to pre-skill state.
Aug 4, 2026
91a1af4
feat(example): azure_doc_qa workflow through SKILL complete.
Aug 5, 2026
c2c11d5
feat(examples): clear seven examples to pre-skill state for rerun.
Aug 5, 2026
c766beb
Add ACS governance for science_research_agent disclosure risks
Aug 5, 2026
20a557e
Add Clarity protocol + ACS governance for change_control_agent
Aug 5, 2026
c6ad69e
Add Clarity protocol + ACS governance for prompt_agents health assistant
Aug 5, 2026
818f7c7
travel_planner_langgraph: Clarity->ASSERT->ACS governance cycle
Aug 5, 2026
fcd361b
feat(example): clear travel_planner_langgraph to pre-skill state.
Aug 5, 2026
07c6776
azure_doc_qa: Clarity protocol for fabrication + leakage risks
Aug 5, 2026
c317619
travel_planner_neurosan: Clarity protocol + ACS governance for two risks
Aug 5, 2026
f838143
azure_doc_qa: grounded fabrication gate + leakage gate through the SKILL
Aug 5, 2026
a1e3b63
feat(example): travel_langgraph_planner ran through SKILL workflow.
Aug 5, 2026
3c94c4c
chore(examples): strip ACS artifacts from the 8 worked domains
Aug 5, 2026
96e3764
feat(examples): rewrite the 8 domain READMEs for the stripped branch.
Aug 5, 2026
c110af7
feat(examples): aligned READMEs to impermissible/permissible behavior…
Aug 5, 2026
a435e4f
feat(readme): update readme with SKILL get started.
Aug 5, 2026
295f687
Merge remote-tracking branch 'origin/main' into ango10/assert-acs-ski…
Aug 5, 2026
4c3a92b
docs(examples): align incident_triage_agent row with main's baseline-…
Aug 5, 2026
58a4eb7
docs: repoint example config paths at the evals/<risk>/ layout
Aug 5, 2026
6e39fa8
feat(example): added the rest of eval_config.yaml for career_health_a…
Aug 6, 2026
81442ff
docs(examples): commit the behavior taxonomy alongside eval_config.
Aug 7, 2026
78db11c
fix(tests): restore-judge preset coverage lost to a silent skip.
Aug 7, 2026
8e6cfc8
fix(tests): collect and trigger the skill's test suite.
Aug 9, 2026
d9be531
fix(deps): bound arize-phoenix below the release that breaks the CI.
Aug 9, 2026
aebdcb6
feat(example): give billing_support_agent a real policy instead of a …
Aug 9, 2026
854f4b8
fix: close the four non-blocking PR review follow-ups.
Aug 10, 2026
76ebb38
fix: close the second-round PR review items.
Aug 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
107 changes: 107 additions & 0 deletions .claude/skills/run-assert-eval/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# run-assert-eval skill

Take a developer from **"I don't know my risks"** to a **measured violation rate
per risk** — without leaving the coding assistant. Risk discovery is owned by
**Clarity** (microsoft/clarity-agent); measurement is owned by **ASSERT**
(responsibleai/ASSERT). This skill wires the two together.

## Files

| File | Purpose |
| --- | --- |
| `SKILL.md` | Claude Code skill entry (the canonical instructions). |
| `../../.github/prompts/run-assert-eval.prompt.md` | GitHub Copilot mirror. |
| `../../.cursor/rules/assert.mdc` | Cursor mirror. |
| `workflows/measure-clarity-failures.md` | The 9-step measurement workflow (parse → triage → configs → run → report → close loop → archive protocol). |
| `workflows/govern-and-remeasure.md` | The ACS governance workflow: turn a measured failure into a deployable ACS policy (`assert-ai acs generate`), wrap the agent, and re-run the same eval to prove the failure rate dropped. |
| `workflows/diagnose-acs-delta.md` | Fallback reference manual for when a governed run's delta comes out wrong (no drop, or over-gating rose) — symptom-indexed, 15 rules. Most are prevented by the pre-flight classification in `govern-and-remeasure.md` Step 1a. |
| `clarity_intake.py` | Dependency-free parser: Clarity failure docs → ASSERT candidate behaviors. |
| `tests/` | Pytest suite + real Clarity fixtures for the parser. |
| `SETUP-CHECKLIST.md` | One-time in-IDE MCP setup + end-to-end verification. |

Keep the three skill surfaces (`SKILL.md`, the Copilot prompt, the Cursor rule)
methodologically aligned when changing the flow.

## Architecture

1. **Discovery (Clarity, shipped):** the Clarity **MCP server** exposes tools —
`run_clarity`, `write_protocol_document`, `record_failure`, `record_suggestion`,
and others. `run_clarity` returns Clarity's real process guide inlined; the host
agent conducts the clarifying conversation and persists findings. See
`SETUP-CHECKLIST.md` to wire it up.
2. **Handoff (files, not JSON):** Clarity writes `.clarity-protocol/`. The
measurement side reads `failures/failures.md` (index) and `failure-NN-*.md`
(individual docs). Those files are the **source of truth**; the parser's JSON is
a disposable cache. Note it is **gitignored, single-domain scratch** — the next
`run_clarity` overwrites it, so each domain's protocol is archived to
`examples/<domain>/Clarity Protocol/` at the end of its run (Step 9), guarded by
a blocking check before any fresh discovery.
3. **Measurement (this skill):** `clarity_intake.py` turns failure docs into
candidate behaviors; `workflows/measure-clarity-failures.md` runs a **mandatory
human triage gate**, generates **one atomic `eval_config.yaml` per selected
failure**, runs them sequentially, and reports one behavior per column.
4. **Governance (ACS, optional):** when a run surfaces a real failure the user wants
to *fix and prove*, `workflows/govern-and-remeasure.md` first **classifies the
failure against the baseline** (Step 1a — semantic `output` gate vs. structural
tool gate, and whether the harm actually routes through the tool being gated),
then derives a deployable **ACS** policy from the findings
(`assert-ai acs generate`), wraps the agent's high-risk tools (or its output),
and re-runs the **same** eval against the governed target to show the
failure-rate delta (baseline → governed). If that delta comes out wrong,
`workflows/diagnose-acs-delta.md` is the symptom-indexed fallback.

## The parser (`clarity_intake.py`)

```
python .claude/skills/run-assert-eval/clarity_intake.py .clarity-protocol
```

Per failure mode it emits a `CandidateBehavior`:
`{name, description, severity, priority, source_doc, candidate_dimensions,
multi_behavior, suggested_splits, warnings}`.

- **Severity → priority**: Critical→P1, High→P2, Medium→P3, Low→P4. Ranges (e.g.
`Medium–Critical`) collapse to the **maximum** severity.
- **Dimensions**: the doc's **Variants** list → an `elicitation_variant` stratify
dimension (highest value — each variant is a distinct route to the failure);
**Failure Chain** conditions → an `interaction_condition` dimension.
- **Atomicity**: docs that bundle several independently testable behaviors are
flagged `multi_behavior` with `suggested_splits` so triage can surface the split.
- **Tolerant**: unknown severity labels or missing headers degrade to a **flagged**
candidate (`warnings` populated) — never a crash, never a silent drop.

Run the tests:

```
python -m pytest .claude/skills/run-assert-eval/tests/test_clarity_intake.py
```

## Worked example

A full end-to-end walkthrough (one P1 — `user_disengagement` — from parse through
triage, config generation, run, headline metrics, and closing the loop) lives in
`workflows/measure-clarity-failures.md` under **Worked example (one P1)**. The ACS
governance counterpart is in `workflows/govern-and-remeasure.md`.

## Related ASSERT docs

Product behavior is documented under `docs/` (team-maintained, on `main`); the skill
**links** rather than restates it — `guides/create-evaluation.md` + `config/schema.md`
(config authoring), `targets/callable.md` (callable signature, return types, OTel
auto-instrumentation) + `targets/model-and-tools.md` (target shapes),
`guides/troubleshooting.md`, `guides/results.md`, `guides/use-local-viewer.md`, and
`guides/securing-agents-with-acs.md` (the ACS loop). This skill owns the *methodology*;
those own *product behavior*. The exceptions the skill documents itself are the two
callable traps those docs omit: `history` is detected by parameter **name** (misnaming it
silently degrades multi-turn to single-turn), and module resolution falls back
`sys.path` → config dir → cwd → direct file load.

## Guarantees the skill enforces

- One atomic behavior per config — never bundle.
- The triage gate and the pre-run confirmation are **human** decisions; declining
writes nothing and runs nothing.
- `.clarity-protocol/` files are authoritative; derived JSON is a cache.
- Discovery goes through Clarity's real MCP tools — no plain-language fallback, no
shelling out to a `clarity cli` process, no separate app.
- Never read/print/commit `.env`, credential values, or `artifacts/`.
79 changes: 79 additions & 0 deletions .claude/skills/run-assert-eval/SETUP-CHECKLIST.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
# Setup checklist — Clarity MCP ⇄ ASSERT (in-IDE only)

These steps require a real IDE with MCP support (VS Code + Copilot agent mode,
Claude Code, or Cursor) and cannot be completed from a headless terminal. Do them
once per workspace, then the `run-assert-eval` skill's discovery front door
(`run_clarity`) becomes callable.

## Phase 1 — Environment setup

- [ ] **Install Clarity with the MCP extra** from your clarity-agent checkout
(Python 3.12+):
```
pip install -e ".[mcp]" # or: uv pip install -e ".[mcp]"
```
- [ ] **Embed Clarity into this project** (generates `.vscode/mcp.json`, the
`.clarity-protocol/` scaffold, and the Clarity-managed block in `AGENTS.md`):
```
clarity embed .
```
- Verify `.vscode/mcp.json` has a `clarity-agent` stdio entry with
`CLARITY_PROJECT_DIR` set to this workspace folder.
- If your Clarity checkout is **uv-managed**, verify the entry uses
`uv run --extra mcp --directory <checkout> python -m clarity_agent.mcp`.
- [ ] **Confirm the server starts**: `python -m clarity_agent.mcp --help`.
- [ ] **Confirm an LLM provider is configured**: `clarity doctor`. Clarity
supports GitHub Copilot, Anthropic, OpenAI, Azure AI, and Gemini. Surface any
failure with its fix — do not silently continue.
- [ ] **Verify ASSERT**: `assert-ai --help`, and a smallest-sample dry run of one
repo example config is invocable.
- [ ] **Reload MCP servers** in the IDE so the `clarity-agent` tools appear, then
confirm you can call `run_clarity`.
- [ ] **Do _not_ commit `.vscode/mcp.json`.** It is generated by `clarity embed .`
and its `uv --directory` arg holds an **absolute, machine-specific path** to
*your* clarity-agent checkout — committing it would hand teammates a broken
path. It is gitignored; each developer runs `clarity embed .` to generate their
own. (If your team pip-installs clarity-agent as a package, the generated entry
is `python -m clarity_agent.mcp` with no absolute path and could be committed —
but the default uv-checkout form must stay local.)

## Phase 2 — End-to-end verification (definition of done)

- [ ] **Fresh discovery**: with no `.clarity-protocol/failures/failures.md`, call
`run_clarity`, conduct a short clarifying conversation, and confirm
`failures.md` gets written.
- [ ] **Parser**: `python .claude/skills/run-assert-eval/clarity_intake.py .clarity-protocol`
emits candidate behaviors; run the unit tests:
```
python -m pytest .claude/skills/run-assert-eval/tests/test_clarity_intake.py
```
- [ ] **Single-P1 run**: from an existing `failures.md`, the workflow presents
triage, you pick one P1, **exactly one** config is generated with a
variants-derived dimension, `assert-ai run` completes, and the results table
renders with the behavior as a column.
- [ ] **Two failures → two configs**: selecting two failures produces two separate
configs and two sequential runs — never one merged config.
- [ ] **Decline at triage → zero writes**: declining at the triage gate results in
zero files written and zero runs.
- [ ] **Loop-close**: a `record_suggestion` round-trip lands in the Clarity mailbox
after a completed run.

## Notes

- Copilot agent mode supports MCP **tools** only (not `clarity://…` resources).
`read_protocol_document` covers the same ground as the resource endpoints.
- Never read/print/commit `.env` or credential values — reference env var **NAMES**
only (AZURE_API_KEY, AZURE_API_BASE, OPENAI_API_KEY, GITHUB_TOKEN,
ANTHROPIC_API_KEY, azure_ad_token).
- Do not edit inside the Clarity-managed block in `AGENTS.md`
(between `<!-- clarity-begin -->` and `<!-- clarity-end -->`).
- **Committing `.clarity-protocol/`**: this repo gitignores it because the protocol
describes a *system-under-test*, not this framework — it's per-target runtime
output. In **your own product's repo**, the protocol describes your product, so
prefer committing the durable docs (`goal/`, `solution/`, `failures/`) and
ignoring only `transcripts/` (and optionally `mailboxes/`). When you finish a
domain here, archive its protocol into `examples/<domain>/Clarity Protocol/` and
**commit it** so it is preserved alongside that domain's `evals/` and `acs/` —
this is Step 9 of `workflows/measure-clarity-failures.md`, and a blocking gate
before any fresh `run_clarity` enforces it (the source dir is gitignored, so an
overwrite is unrecoverable). See the per-example replication package in `SKILL.md`.
Loading
Loading