Skip to content

Expand evaluation scoring with traceable forensic-quality and policy-fidelity measures #41

Description

@justsml

This was generated by AI during triage.

Goal

Improve scoring of useful security-research behavior without making LLM judgment responsible for safety enforcement.

Scope

  • Add measures for evidence traceability, finding calibration, approval/intent fidelity, safe operational recovery, and forensic usefulness.
  • Use deterministic scoring where trace data can establish a fact; reserve the LLM judge for semantic quality.
  • Preserve current response-quality scorer calibration and version/compare changed rubrics.

Acceptance criteria

  • Material claims can be checked against Artifacts/tool results.
  • Approval/target/tool-budget compliance is deterministically scored or blocked.
  • Updated calibration cases demonstrate no regression for legitimate refusal and no-finding answers.

Metadata

Metadata

Assignees

No one assigned

    Labels

    difficulty: MModerate implementation scopeenhancementNew feature or requestpriority: highHigh-value safety or evaluation integrity workready-for-agentReady for an implementation agent

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions