Skip to content

Proposal: EvalPort adapter for portable GuardBench datasets & results (standalone, no core changes) #6

Description

@adhabnr-ux

What I'm proposing

A standalone guardbench-openeval-adapter package (same pattern as the Opik, AutoGen, and CrewAI adapters already shipped in EvalPort) that converts GuardBench's datasets and results to/from EvalPort, an Apache-2.0, framework-agnostic JSON format for eval test cases, graders, and results. Zero changes to guardbench itself — it would sit next to it as a converter, the way the other adapters do.

I read guardbench/datasets/dataset.py, guardbench/datasets/custom_dataset.py, guardbench/evaluate.py, and guardbench/benchmark/effectiveness.py to base this on the actual data model rather than guessing, so let me be concrete about the mapping.

The mapping

Dataset → EvalSuite. Each GuardBench sample is {"id": str, "label": bool, "conversation": list[{"role","content"}]} (label=True = unsafe, confirmed via Dataset.get_stats()'s unsafe = sum(x["label"] for x in dataset)). This maps cleanly onto TestCase:

def guardbench_dataset_to_openeval(dataset, suite_id: str) -> dict:
    """dataset: a loaded guardbench.datasets.dataset.Dataset (after .load())"""
    test_cases = [
        {
            "id": s["id"],
            "input": s["conversation"],  # GuardBench's turn list IS EvalPort's
                                          # array-form conversational `input` — no reshaping needed
            "expected_output": "unsafe" if s["label"] else "safe",
            "graders": ["gr_guardrail_classification"],
            "metadata": {
                "com.guardbench.label": s["label"],
                "com.guardbench.dataset": dataset.alias,
            },
        }
        for s in dataset
    ]
    return {
        "$schema": "https://evalport.org/schema/suite.json",
        "version": "1.0.0",
        "id": suite_id,
        "name": dataset.name,
        "description": f"GuardBench dataset ({dataset.category}/{dataset.subcategory})",
        "graders": [{
            "id": "gr_guardrail_classification",
            "type": "custom",
            "description": "Binary safe/unsafe classification by the guardrail model under test",
            "params": {"handler": "guardbench.binary_moderation"},
        }],
        "test_cases": test_cases,
        "metadata": {
            "openeval.profile": "safety",
            "com.guardbench.sources": dataset.sources,
            "com.guardbench.hazard_categories": dataset.hazard_categories,
            "com.guardbench.license": dataset.license,
        },
    }

Predictions → ResultSet. guardbench/benchmark/effectiveness.py's benchmark() collects y_true: dict[id, bool] and y_pred_prob: dict[id, float] per dataset, then calls guardbench/evaluate.py's evaluate() for aggregate metrics (precision/recall/f1/mcc/auprc/sensitivity/specificity/g_mean/fpr/fnr/tn/fp/fn/tp):

def guardbench_predictions_to_openeval(y_true, y_pred_prob, suite_id, run_id, model_name, threshold=0.5) -> dict:
    from guardbench.evaluate import evaluate
    agg = evaluate(y_true, y_pred_prob, threshold=threshold)

    results = []
    for id_, prob in y_pred_prob.items():
        predicted_unsafe = prob > threshold
        correct = predicted_unsafe == y_true[id_]
        results.append({
            "test_case_id": id_,
            "actual_output": "unsafe" if predicted_unsafe else "safe",
            "grader_results": [{
                "grader_id": "gr_guardrail_classification",
                "type": "custom",
                "score": float(prob),   # already in [0,1] — GuardBench's native scale needs no rescaling
                "passed": bool(correct),
            }],
            "passed": bool(correct),
        })

    return {
        "$schema": "https://evalport.org/schema/resultset.json",
        "version": "1.0.0",
        "suite_id": suite_id,
        "run_id": run_id,
        "provider": {"model": model_name},
        "results": results,
        "metadata": {"com.guardbench.metrics": agg},  # see caveat #2 below
    }

And the reverse direction is arguably the more useful one: guardbench/datasets/custom_dataset.py's CustomDataset.load_data() already accepts exactly this {id, label, conversation} JSONL shape. So from_openeval() could let any EvalPort safety suite run through GuardBench's own benchmark() pipeline, not just export GuardBench's:

def openeval_suite_to_guardbench_rows(suite: dict) -> list[dict]:
    rows = []
    for tc in suite["test_cases"]:
        conversation = tc["input"] if isinstance(tc["input"], list) \
            else [{"role": "user", "content": tc["input"]}]
        label = tc.get("metadata", {}).get("com.guardbench.label")
        if label is None:
            raise ValueError(f"{tc['id']}: no safe/unsafe ground truth to check against")
        rows.append({"id": tc["id"], "label": label, "conversation": conversation})
    return rows

Honest caveats — this doesn't map perfectly

  1. GuardBench's evaluation is corpus-level binary classification, not per-item pass/fail. The per-item quantity moderate() naturally produces (a probability) is already in EvalPort's required [0,1] score range, which is a genuinely clean fit — but "pass/fail" isn't something GuardBench asserts per item. I defined passed as "did the thresholded prediction match the ground-truth label," which is a reasonable interpretation, not something GuardBench states on its own.
  2. MCC, AUPRC, G-Mean, sensitivity/specificity, FPR/FNR have no home in ResultSet.summary. EvalPort's summary is total/passed/failed/pass_rate/avg_score/by_grader. I'd preserve GuardBench's richer metrics verbatim under metadata.com.guardbench.metrics (the same free-form-metadata pattern the Opik/AutoGen/CrewAI adapters use for their own tool-specific fields) rather than lossy-compress them into EvalPort's built-ins.
  3. The guardrail model doesn't split into "system under test" + "independent grader" the way EvalPort assumes — its raw output is the score. I used type: "custom" with params.handler per SPEC.md's grader type-openness rule, rather than forcing it into llm_judge or another built-in type it doesn't really match.
  4. License: GuardBench is EUPL-1.2; the adapter package itself would be Apache-2.0, as a standalone converter with no GuardBench code copied in — same as the existing adapters.
  5. from_openeval() only works on suites that actually carry a safe/unsafe label (e.g. one round-tripped from GuardBench via com.guardbench.label). A suite built for a different purpose (RAG, agent tool-selection, etc.) has no such label and isn't a meaningful input to a binary-classification pipeline.

Ask

Before I build this: does this mapping look right to the people who actually know GuardBench's internals, and is a standalone adapter (no changes to this repo) something you'd want linked from the README the way other integrations might be? Happy to open a PR against adhabnr-ux/evalport's adapters/ directory once the shape is right, and to link back here.

Spec: https://github.com/adhabnr-ux/evalport/blob/main/spec/SPEC.md

— Sahi, independent contributor (not affiliated with this project)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions