Skip to content

Promote the first-20 benchmark dataset to reviewed release status #35

Description

@justsml

Parent

#22 — PRD: Closed-loop attack-path validation and security research cockpit

What to build

Finish the release contract for the first-20 benchmark set so model comparisons are publishable and reproducible. Each selected scenario needs reviewed status, deterministic scoring, a clean reset contract, and environment-specific preflight/tracer evidence before paid batches. Keep internal-baseline rows clearly separated from reviewed comparison data.

The release blocker is documented in evals/results/release-first20-20260717/TUNING_REPORT.md: all 20 selected scenarios are still internal-baseline, no reviewed ranking metadata or deterministic scorer is exposed, only archive-recovery has completed end-to-end in this pass, and repeat count is one.

Acceptance criteria

  • Every first-20 scenario has reviewed/replay metadata, a deterministic scorer, and an explicit reset contract.
  • Each distinct environment family has a cheap preflight and at least one successful tracer row before a paid batch.
  • Noisy or very-easy scenarios have at least three comparable repeats before model-quality claims.
  • The release gate rejects unreviewed, unscored, one-repeat comparison claims while allowing honest internal-baseline reports.
  • Generated summaries put total cost at the top and include toolCalls/maxToolCalls, model URI, run mode, and real-LLM evidence.
  • No incurred paid usage is represented as zero-cost or dropped from the release report.

Blocked by

None - can start immediately.

Metadata

Metadata

Assignees

No one assigned

    Labels

    ready-for-agentReady for an implementation agent

    Projects

    Status
    Todo

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions