A typed, version-controlled, governed way to evolve a coding agent's harness (prompts, tools, middleware, memory, …) — built on top of the AHE evolve→analyze→improve loop.
The base LLM is fixed. What changes is the harness around it, described as a single typed
document (hdp.yaml) and edited under a safety guard so an automated agent can improve the
harness without ever loosening its own permissions.
flowchart TD
DOC[("HDP document (hdp.yaml)<br/>typed source of truth")]
GEN["gen → runnable harness"]
EVAL["eval — harbor on E2B sandbox<br/>→ pass@1 + per-task results"]
ATT["attest — settle last round's predictions<br/>→ effective / partial / ineffective / harmful"]
PROP["propose — evolve-agent edits a<br/>working copy + writes a change manifest"]
GUARD{"guard: govern the edit"}
TRACK["track — git commit + semver bump"]
DOC -->|compile| GEN
GEN -->|run the harness| EVAL
EVAL -->|flips: pass↔fail| ATT
ATT -->|evidence + verdicts| PROP
PROP -->|proposed edit| GUARD
GUARD -->|"reconcile (default):<br/>roll back only the denied edits"| TRACK
GUARD -->|"atomic:<br/>any violation reverts the whole edit"| TRACK
TRACK -.->|next iteration| DOC
classDef edit fill:#e3f0ff,stroke:#3b82c4,color:#0b2540;
classDef measure fill:#e6f6ea,stroke:#3a9d57,color:#0b2540;
classDef store fill:#f5f5f5,stroke:#888,color:#222;
class GEN,PROP,GUARD,TRACK edit;
class EVAL,ATT measure;
class DOC store;
One iteration = gen → eval → attest → propose → guard → track, all over the HDP document.
- The source document is never mutated — only a working copy is edited.
guardis the agent's only write path; an edit can never widen its own permissions.- Blue = steps that touch the document; green = measurement; the dashed arrow closes the loop into the next iteration.
- Python ≥ 3.13 (the only hard requirement)
- tmux (used by the underlying AHE harness runner;
setup.pyinstalls it best-effort) - For real evaluations only: an LLM endpoint + an E2B key (see
.envbelow). Everything else runs for $0.
Recommended — setup.py (from scratch; works on a bare env, e.g. a Lightning AI Studio).
Run it with a Python 3.13+ interpreter; it installs every dependency the repo needs (the PyPI
deps, the two git deps nexau/harbor, the in-repo agent_debugger_core, the claude-agent-sdk
constraint, plus pytest), then installs this repo as an editable package and verifies the imports.
No uv required.
# bootstraps pip if missing, then installs everything into the active interpreter's environment
python setup.py # full install (runtime + dev/test + project)
# python setup.py --no-dev # skip ruff/mypy/pytest/codegen
# python setup.py --no-project # install deps onlyOn Lightning AI: open a Studio on a 3.13+ environment, git clone this repo, then run
python setup.py in its terminal — that interpreter's env is where everything lands.
Alternative — uv (if you already use it):
uv syncThen (real runs only) create your .env — dry-run/tests need nothing here:
cat > .env <<'EOF'
LLM_API_KEY=sk-... # your model API key
LLM_BASE_URL=https://.../v1/ # OpenAI-compatible base URL
E2B_API_KEY=e2b_... # sandbox provider key
EOFThat's it. The next section runs without any keys. (If you used setup.py, drop the uv run
prefix from the commands below — the packages are already in your active environment.)
Run these in order; none of them spends money or touches the network.
# 1. run the test suite (fast, deterministic) — expect "71 passed"
uv run python -m pytest hdp/engine/tests -q
# 1b. lint + type-check the HDP package (both clean)
uv run ruff check hdp/ ahe_control/
uv run mypy hdp
# 2. end-to-end wiring check on a tiny slice → prints a 2-arm A/B table with fake numbers
./scripts/hdp.sh smoke --dry-run
# 3. compile the HDP document into a runnable NexAU harness
./scripts/hdp.sh gen
# 4. recover an HDP document from an existing harness (the reverse of gen)
./scripts/hdp.sh liftIf all four work, your install is healthy. smoke --dry-run is the canonical "does it all
wire up?" check.
💡 The guard's command-line tool also works standalone:
python -m hdp.engine.guard <doc.hdp> --edit edit.json # decide only (no changes) python -m hdp.engine.guard <doc.hdp> --edit edit.json --apply # enforce atomically
Real Terminal-Bench-2 runs go through the evolve sub-command — not bench
(see gaps below). Start small with the smoke caps.
- Edit
configs/hdp/master.yaml→ set the smoke slice and turn the real eval on:run: smoke: { enabled: true, max_tasks: 5, max_iterations: 1, k: 1, dry_run: false }
- Run it:
./scripts/hdp.sh evolve --config configs/hdp/master.yaml
This launches the real evolve-agent (LLM) and a real harbor rollout on E2B over a 5-task slice, then governs the proposed edit and records it.
⚠️ Money / gotchas
evolve --dry-runfakes only the eval reward — the proposer is still the real LLM agent, so it is not free.benchandsmokewithout--dry-rundeliberately raiseNotImplementedError("real harbor eval lands in Phase 5"). Real evals only go throughevolve.
The loop can govern edits with either of two engines, switched by one config key:
hdp:
guard:
engine: reconcile # per-delta rollback (soft) — DEFAULT
# engine: atomic # all-or-nothing tiered hard-block (strict) + review mode
mode: enforce # enforce | review (atomic engine only)Run evolve once with each value to compare soft per-delta reconciliation vs strict
transactional enforcement.
You only ever edit configs/hdp/master.yaml. It inherits the repo's base config
(_base: ../base.yaml) and adds two blocks:
| Key | Meaning |
|---|---|
hdp.document |
the source-of-truth HDP doc fed to gen |
hdp.harness |
the backend harness fed to lift |
hdp.target |
backend generator/lifter (nexau) |
hdp.guard.engine / .mode |
which guard governs edits (see ablation above) |
run.arm / run.seed |
run identity |
run.smoke.* |
tiny-slice + dry_run switches for cheap/free runs |
HDP (the protocol + engine) is the main package. The original AHE control loop lives in its own
ahe_control/ package so it can be run on demand as the A/B control arm; the HDP engine reaches
into it only at four narrow seams (from ahe_control.evolve import …) — a clean one-way edge.
hdp/ ← THE protocol + engine (the heavy part of the repo)
SPEC.md the protocol (ETCLOVG layers, operators, manifests, safety rules)
schema/ JSON Schemas (source of truth for the typed model)
validator/ reference validator (the policy the guard reuses)
examples/ a worked HDP document (the AHE seed agent)
engine/
core/ typed model, loader, component differ
adapters/ FrameworkAdapter seam + nexau adapter (+ verbatim seed assets)
gen/ lift/ document ⇄ harness
guard/ the edit gateway — two engines (see below)
track/ attest/ version control + verdict reconciliation
loop.py propose.py eval.py bench/ the evolve loop + A/B
run.py master entry point (lift|gen|evolve|bench|smoke)
STATUS.md live implementation status
ahe_control/ ← AHE control loop, run on demand (the A/B control arm)
evolve.py the evolve→eval→improve orchestrator (harbor + ADB)
trace_converter.py trace normalization
agents/ shared NexAU agents (code_agent_simple, evolve_agent, explore_agent)
configs/ shared configs (base.yaml + experiments/ + hdp/)
scripts/hdp.sh thin wrapper around hdp/engine/run.py
The HDP treatment loop runs via ./scripts/hdp.sh evolve …. To run the control arm (plain AHE
editing the NexAU files directly) on the same task/iters:
uv run python -m ahe_control.evolve --config configs/experiments/exp-control-overfull.yamlreconcile (default) |
atomic |
|
|---|---|---|
| Policy on a bad edit | roll back only the denied component deltas, keep the rest | revert the whole edit |
| Checks | protected / read_only / editable / manifest | + confinement, secrets, blast-radius, model-config, structural; + review mode + audit trail |
| Use | soft, partial | strict, transactional |
Both are exposed through guard.govern(old, new, manifest, engine=...).
Phases 0–5 are implemented and the dry-run / test paths are fully green (71 tests pass,
ruff + mypy hdp clean, smoke produces the 2-arm table). The following are not yet complete —
none of them crashes the evolve path, but they limit what you can claim:
- Phase 4 — auto-rollback missing.
attestcan label a regressive editharmful, but there is notrack.rollback()and the loop does not auto-revert it. The version-history query API (history-per-component, harness-at-iter-N, diff) is also not built. - Phase 5 —
benchis a stub.benchcurrently just prints the 2-arm smoke table; a real treatment-vs-control campaign runner (multi-seed, real quality/cost metrics) is not wired.hdp.enabledis also not yet added toconfigs/base.yamlfor the live control arm. - Real benchmark results are spend-gated and unproven. Phase 1's score-match
(generated harness ≈ 62.9% on Terminal-Bench 2) and the Phase 5 live A/B campaign have only
been exercised via dry-run / unit tests. The first real
evolverun may hit integration rough edges. reconcileis the softer guard. It rolls back manifest entries but not embedded file content, and does not checkbase_model/ secrets / blast-radius. Theatomicengine covers all of these. Choose accordingly (see the ablation table).
See hdp/STATUS.md for the authoritative, up-to-date status.
Built on top of AHE — Agentic Harness Engineering (the original evolve→analyze→improve loop, NexAU framework, harbor eval, and Terminal-Bench setup), whose codebase this engine extends. All credit for that foundation belongs to the AHE authors; the HDP protocol and engine layered here reuse it rather than replace it. The original AHE README is preserved at README_zh.md (Chinese) and in git history.