Mendmark 0.4 is ready for a small design-partner pilot with teams shipping tool-using agents.
The useful outcome is concrete: run Mendmark against an existing eval suite, inspect any surviving controlled faults, strengthen an evaluator if warranted, and measure whether the audit belongs in CI.
Good pilot fits:
- At least one passing eval case for a tool-using agent
- An engineer who can update the evaluator or release policy
- DeepEval or the ability to exchange local JSON with an evaluator command
- Permission to share aggregate setup time, runtime, mutation counts, and survivor counts
Mendmark runs locally. Do not post prompts, traces, arguments, outputs, credentials, or customer data here.
Read the pilot guide and apply through the repository's Mendmark pilot request issue form. The first cohort will be limited to 5–10 teams so onboarding can remain hands-on.
Mendmark 0.4 is ready for a small design-partner pilot with teams shipping tool-using agents.
The useful outcome is concrete: run Mendmark against an existing eval suite, inspect any surviving controlled faults, strengthen an evaluator if warranted, and measure whether the audit belongs in CI.
Good pilot fits:
Mendmark runs locally. Do not post prompts, traces, arguments, outputs, credentials, or customer data here.
Read the pilot guide and apply through the repository's Mendmark pilot request issue form. The first cohort will be limited to 5–10 teams so onboarding can remain hands-on.