Skip to content

jramos/agent-self-evolution

 
 

Repository files navigation

🧬 Agent Self-Evolution

tests

Automated evolution for the skills, prompts, and tool code your agent runs on.

Agent Self-Evolution proposes better versions of the artifacts your agent depends on — skills, tool descriptions, system prompts, tool implementations — and ships only the ones that prove out against the real agent. A noise-aware deploy gate, a closed-loop behavioral suite, and an executable-test oracle for code separate genuine gains from lucky ones, so your agent improves without ever shipping a regression. No GPU, no training — everything runs via API calls at ~$1–5 per run.

Works on any agent framework that emits SKILL.md filesHermes Agent is the original target; Claude Code and any other <dir>/<skill>/SKILL.md layout are supported via a pluggable skill-source abstraction.

How It Works

flowchart LR
    A[Read current<br/>skill/prompt/tool] --> B[Generate<br/>eval dataset]
    B --> C[GEPA<br/>Optimizer]
    C --> D[Candidate<br/>variants]
    D --> E1[Synthetic<br/>holdout]
    D --> E2[Closed-loop<br/>behavioral suite]
    E1 -. Execution traces .-> C
    E1 --> F["Dual-signal deploy gate<br/>(synthetic + closed-loop;<br/>CL-primary on synth-tie)"]
    E2 --> F
    F --> G[Best<br/>variant]
    G --> H[PR against<br/>source repo]
Loading

GEPA reads execution traces to understand why things fail, then proposes targeted improvements (ICLR 2026 Oral, MIT). On top of GEPA the framework adds a held-out deploy check, three-dimensional judge scoring, and a closed-loop behavioral suite, so a candidate that just got lucky on a small eval set can't ship — the argument for why these layers matter is in docs/framework_advantages.md. A saturation pre-flight refuses to spend budget on a baseline with no headroom, and closed-loop validation runs the real agent on a task suite to confirm behavior actually shifted (both in the usage guide).

Quick Start

git clone https://github.com/jramos/agent-self-evolution.git
cd agent-self-evolution
uv sync

Evolve a skill with Hermes Agent. Running Hermes? No env vars — it auto-detects ~/.hermes/config.yaml for provider, model, and credentials.

uv run python -m evolution.skills.evolve_skill --skill github-code-review --iterations 10

Evolve a skill with any provider. Set a standard env var and run:

export ANTHROPIC_API_KEY=sk-ant-...
uv run python -m evolution.skills.evolve_skill --skill writing-skills --iterations 10

Evolve a Claude Code CLAUDE.md convention — the clearest win: a project-specific convention the base prompt can't know (e.g. "run tests with ./bin/check, never pytest"), driven headlessly and scored by convention adherence (no LLM judge):

uv run python -m evolution.prompts.evolve_prompt_section \
    --target claude --section REPO_CONVENTIONS --claude-md ./CLAUDE.md \
    --tasks evolution/validation/suites/claude_conventions.jsonl --agent-model sonnet --apply

Full usage guide for every target (tools, prompts, session-history evals, tuning, shipping) → docs/usage.md · all CLI flags → docs/interfaces.md · or run --help on any command.

What You Can Evolve

Phase Target Engine Status
Phase 1 Skill files (SKILL.md) DSPy + GEPA Mechanism validated
Phase 2 Tool descriptions + dual-signal deploy gate DSPy + GEPA Mechanism validated
Phase 3 System prompt sections (Hermes + Claude Code) DSPy + GEPA Mechanism validated
Phase 4 Tool implementation code Iterative test-feedback repair Validated
Phase 5 Continuous improvement loop Propose-only triage sentinel Sentinel shipped

Phases 1–3 are validated as a working mechanism (the pipeline runs end-to-end and the gate catches regressions); on a capable agent, evolving these artifacts mostly catches regressions rather than finding improvements — see Findings for where evolution actually pays off.

Run Phases Together

Sequence the per-phase evolvers from one run-spec instead of invoking each by hand. The orchestrator runs them in order (skills → tools → prompts → code), isolates each as a subprocess, and writes a JSONL run history + summary to its run root. It is propose-only: it never deploys, and opens no PRs unless you pass --allow-pr.

# Validate a run-spec without spending (prints each phase's resolved command):
python -m evolution.orchestrator --spec examples/orchestrator/sample_run.yaml --dry-run

# Run it for real — continue-on-error by default; --only restricts to a subset:
python -m evolution.orchestrator --spec examples/orchestrator/sample_run.yaml --only skills --only tools

See examples/orchestrator/sample_run.yaml for the spec format.

Use the Gate Without Evolution

The gate is useful whether or not GEPA generated the candidate. Bring your own change and ask whether it's real — none of these run evolutionary search:

# "Did my hand-written tool-description change actually help the real agent?"
# Real-agent A/B with an A/A noise floor, so a within-noise gain can't deploy.
python -m evolution.validation.closed_loop --tool patch --hermes-repo ~/.hermes/hermes-agent \
    --tasks suite.jsonl --baseline baseline.py --evolved my_change.py --noise-aware-gate

# "Repair this broken tool from its failing test — and prove the fix isn't gamed."
# Throwaway worktree + isolated venv; held-out split, surface freeze, file scope, regression floor.
python -m evolution.code.evolve_code --repo ~/.hermes/hermes-agent \
    --tool tools/foo.py --visible-test tests/tools/test_foo_a.py --holdout-test tests/tools/test_foo_b.py

# "What real bugs in this repo's git stream could the loop fix?" ($0, pure git, no LLM)
python -m evolution.monitor --repo ~/.hermes/hermes-agent

The deploy gate (held-out split, surface freeze, baseline-diff regression floor) is the most reusable thing here: point it at any LLM-authored patch and it resists the specific ways a green test lies.

Findings

Skill / prompt / tool-description text changes a capable agent's behavior only where it carries a signal the model can't already infer. The campaign resolves that into three regimes:

  • GREEN — non-inferable signal (proven). An operating convention or project wrapper the agent can't guess: a CLAUDE.md region drove adherence to a counter-default wrapper from 0% → 100% held-out (judge-free oracle, n=10); code repair fixes ~60–74% of real bugs from a failing test. This is the product — install the missing signal; the gate certifies it didn't regress.
  • NULL — binary-correctness reasoning (a measured cliff). Where a task is pass/fail and the agent is already competent, there's no room: capable agents bracketed the headroom band from both sides (aced debugging at 1.00; uniformly failed a 13-requirement conjunction at 0.00). The pre-flight detects this and refuses the run.
  • FRONTIER — open-ended generative quality (open). The regime where "a better prompt yields better output" is most plausible (writing, analysis, review) resists measurement: the LLM judge saturates, and an execution oracle needs a targeted task while open-ended quality is intrinsically un-targeted (reports/frontier_quality_findings.md).

Two consequences worth stating plainly. Evolution earns no magic on its own — on this corpus GEPA is matched by plain best-of-N resampling, and a population-based evolutionary search (the actual imbue darwinian_evolver) was tested at equal budget and strictly dominated by best-of-N (6/23 vs 8/23, recovering nothing best-of-N missed), so the value is the signal + the gate, not the search. And self-evolution reached deploy-grade traction only under a conjunction — an executable oracle, real headroom, and code repair from failing-test feedback (deploy-reachable ~0.60–0.74, clearing a pre-registered futility floor) — which a leakage check confirms is test-feedback repair, not autonomous re-derivation.

Full evidence, CIs, and validity threats: headroom experiments · consolidated findings (PDF) · frontier-quality · evolver vs best-of-N.

Engines

Engine What it does Targets License
DSPy + GEPA Reflective prompt evolution from execution traces skills, tool descriptions, prompt sections MIT
Iterative test-feedback repair (DSPy) Whole-file repair from a failing test in a throwaway worktree, behind the held-out + surface-freeze gate tool implementation code MIT

Guardrails

Every evolved variant must pass: (1) the full test suite (pytest tests/ -q, 100%); (2) size limits (skills ≤15KB, tool descriptions ≤500 chars); (3) caching compatibility (no mid-conversation changes); (4) semantic preservation (no drift from the original purpose); and (5) human PR review — never a direct commit.

Operating the Sentinel

The code-evolution loop has a propose-only front-end: a triage sentinel that scans a target repo's recent git stream for bugs the repair loop could fix, ranks them, and writes a queue. It never evolves code or opens a PR — a human reads the queue and decides what to attempt.

# Scan ($0, pure git, no LLM) — safe to schedule
python -m evolution.monitor --repo /path/to/target-repo --since-days 90

# Attempt the top candidates (the only step that spends; cost-capped, human-gated)
python -m evolution.monitor --repo /path/to/target-repo --attempt-top 3 --max-cost-usd 5.0

See docs/operating_the_sentinel.md for the queue format, the verdict taxonomy, and an opt-in scheduled scan.

Docs & Reference

License

MIT — © 2026 jramos and Nous Research

About

⚒ Evolutionary self-improvement for Agents — optimize skills, prompts, and code using DSPy + GEPA

Resources

Stars

10 stars

Watchers

0 watching

Forks

Releases

No releases published

Packages

 
 
 

Contributors

Languages

  • Python 99.6%
  • Shell 0.4%