Skip to content

benchmark: public competitive evidence and qualification release gate #40

Description

@luke-n-alpha

Goal

Produce independently defensible, multilingual competitive evidence for naia-memory and keep publication fail-closed until the exact benchmark artifacts, authorization chain, deployment trust anchor, and signed qualification can be reproduced and verified.

Public-ready acceptance criteria

  • Competitive claim is based on independently authored/native-reviewed held-out data, not generated diagnostics alone.
  • Naia, Mem0, Graphiti, Hindsight, and any other declared comparator run under a frozen, engine-neutral contract with equivalent model/runtime disclosure.
  • Korean and multilingual results disclose per-language relevance, contradiction/forbidden-result rates, latency, cost, and uncertainty.
  • Dataset custody, preregistration, participant delivery, execution authorization, blind adjudication, and non-peeking constraints verify end to end.
  • A release-owned public deployment policy provisions a static Ed25519 trust anchor without committing private keys.
  • A signed competitive qualification binds all required artifacts and passes the canonical public gate from a clean checkout.
  • Independent adversarial review examines benchmark overfitting, Naia-specific metric advantage, comparator parity, and claim scope.
  • Final Korean/English public report separates demonstrated results, limitations, and unproven claims.

Existing evidence and history

Earlier benchmark/qualification progress was accidentally posted to #39, whose actual scope is conflict-aware structured facts. Those comments remain immutable historical records but are not treated as the canonical tracker:

Relevant pushed commits include 094ba5d, 7392980, a9f00b9, and 6ce69b1.

Current status (2026-08-24)

The public gate and provisioning machinery are materially stronger, but the repository remains intentionally not public-qualified: the committed deployment anchor is null, no release-owned policy/key has been provisioned, and the final independent multilingual comparator campaign has not been sealed against that exact trust-store digest.

GPU1 is out of scope and must not be used.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions