This was generated by AI during triage.
Goal
Improve scoring of useful security-research behavior without making LLM judgment responsible for safety enforcement.
Scope
- Add measures for evidence traceability, finding calibration, approval/intent fidelity, safe operational recovery, and forensic usefulness.
- Use deterministic scoring where trace data can establish a fact; reserve the LLM judge for semantic quality.
- Preserve current response-quality scorer calibration and version/compare changed rubrics.
Acceptance criteria
- Material claims can be checked against Artifacts/tool results.
- Approval/target/tool-budget compliance is deterministically scored or blocked.
- Updated calibration cases demonstrate no regression for legitimate refusal and no-finding answers.
Goal
Improve scoring of useful security-research behavior without making LLM judgment responsible for safety enforcement.
Scope
Acceptance criteria