aorta agent mitigateThis guide explains how agentic testing works in AORTA: what the
aorta agent mitigate command does, whether it uses a real LLM, what happens under
the hood, and how to read the output.
For the high-level design rationale, see aorta-probe-agent.md. For probe-mode mechanics (recipes, classifiers, artifacts), see probe-188/usage.md.
Command name.
aorta agentis now a namespace:aorta agent <name>dispatches to one of the agents registered under theaorta.agentsentry-point group, and the closed-loop mitigation search described here isaorta agent mitigate. The old bare form,aorta agent -- <command>, still works for one release; it prints a deprecation notice on stderr (so stdout stays parseable) and runs the same command. Seesrc/aorta/registry/README.mdfor how to register another agent.
Agentic testing means a closed loop instead of a one-shot run:
Today, aorta probe does step 1–2 across a fixed matrix you write in
YAML. aorta agent mitigate automates steps 3–5 on top of the same engine.
The loop is agentic because it maintains state, makes sequential decisions, and adapts the next experiment from prior results — even when no external LLM is involved.
By default: no.
| Setting | LLM used? | How decisions are made |
|---|---|---|
Default (--llm-backend fake) |
No | Deterministic FakeLLMProposer: heuristics on detector IDs + round-robin through registered mitigations |
--llm-backend litellm |
Yes | LiteLLM calls your configured model; requires pip install 'amd-aorta[agent]' and provider API keys |
--llm-backend vllm / openai |
Yes | The model aorta chat is configured with (~/.config/aorta/chat.toml or AORTA_CHAT_*); requires pip install 'amd-aorta[chat-cli]' |
The CLI default is fake so tests, CI, and local smoke runs work with
zero API calls and fully reproducible behavior.
The LLM runs at one stage only: the proposer step, after probe has
already executed a cell and the deterministic 5-tier classifier has
written result.json.
sequenceDiagram
participant CLI as aorta agent mitigate
participant Loop as agent loop
participant Probe as run_recipe / probe
participant Classifier as 5-tier classifier
participant Proposer as fake or LiteLLM
CLI->>Loop: argv + ticket + policy
Loop->>Probe: run none-none cell
Probe->>Classifier: stdout/stderr/exit/hang signals
Classifier->>Probe: verdict pass/fail + detectors
Probe->>Loop: result.json
alt non-baseline cell passed
Loop->>CLI: outcome converged
else still searching
Loop->>Proposer: cell summaries + candidates + tried list
Proposer->>Loop: category, hypothesis, next_mitigations, stop
Loop->>Probe: grow mitigation_axis, run next cell
end
The LLM never:
pass / fail (only the classifier does).The LLM only:
rccl_hang, illegal_mem, …).Install optional LLM support:
pip install 'amd-aorta[agent]'
export OPENAI_API_KEY=... # or other provider LiteLLM supports
Then:
aorta agent mitigate --llm-backend litellm --llm-model gpt-4o-mini ...
To use a model you serve yourself on vLLM or TokenSpeed, point the chat
settings at it and select the vllm backend:
pip install 'amd-aorta[chat-cli]'
export AORTA_CHAT_LLM_PROVIDER=vllm
export AORTA_CHAT_VLLM_BASE_URL=http://localhost:8000/v1
export AORTA_CHAT_VLLM_MODEL=Qwen/Qwen3-8B
aorta agent mitigate --llm-backend vllm ...
For a Qwen3-family model, start the engine with the flags in serving a Qwen3-family model.
The fake backend still implements a full agent loop:
| Agent property | Implementation |
|---|---|
| Perception | Reads result.json per cell: verdict, failure_detectors_fired, capture |
| Memory | agent_log.jsonl + on-disk probe cells; wake() resumes after crash |
| Planning | Infers category from detector IDs (e.g. tier2:* → rccl_hang); picks next untried mitigation from registry order |
| Action | Appends mitigation to axis, calls run_recipe with flat_resume |
| Termination | Stops on baseline pass, converged mitigation, budget, or exhausted candidates |
| Guardrails | AgentPolicy: max iterations, wall time, registry-only names, optional approval gate |
So “agentic” here means autonomous search over a mitigation space, not “must call Claude/GPT.” The LLM is an optional upgrade for smarter mitigation ordering and richer hypotheses — not a requirement.
# From repo root with src on PYTHONPATH, or after pip install -e .
PYTHONPATH=src aorta agent mitigate \
--output ./agent_results \
--ticket ROCM-EXAMPLE \
-- \
python3 my_repro.py --steps 100
Required: literal -- before your command (same rule as aorta probe).
Useful flags:
| Flag | Purpose |
|---|---|
--symptom "..." |
Hint for proposer (fake or LLM) |
--max-iterations N |
Cap mitigation proposals (default 8) |
--mitigation NAME |
Restrict search (repeatable) |
--mitigations-file sidecar.json |
Extra registered mitigations |
--llm-backend litellm |
Enable real LLM proposer |
--prompt-profile rl-episode |
Send the prompt an RL-trained checkpoint learned on (prompt profiles) |
--dry-run |
Plan cells without executing |
--bundle |
Run aorta bundle after loop (needs recipe redaction) |
-v / -vv |
Progress logging |
Command:
PYTHONPATH=src aorta agent mitigate \
--output /tmp/agent_out \
--ticket smoke-hello \
-- \
echo hello
What happens under the hood:
mitigation_axis: [none].run_recipe runs cell none-none → executes echo hello.verdict: pass.stop_reason: baseline_pass.agent_report.md and stops (no mitigation search).Expected CLI output:
Agent outcome: baseline_pass — Baseline passed — no mitigation search needed.
Wrote /tmp/agent_out/smoke-hello/agent_report.md
Baseline cell (none-none) passed. The repro succeeds without mitigations; no search was run.
Key artifacts:
/tmp/agent_out/smoke-hello/
agent_log.jsonl
agent_report.md
none-none/trial_0/result.json # verdict: pass
matrix.json
host_env.json
Command:
PYTHONPATH=src aorta agent mitigate \
--output /tmp/agent_out \
--ticket smoke-fail \
--max-iterations 3 \
--mitigation none \
--mitigation tf32_off \
--mitigation xnack \
-- \
python3 -c 'import sys; sys.exit(1)'
What happens under the hood:
none-none → exit 1 → verdict: fail,
detectors include tier1:exit_nonzero.launch_error or unknown; proposes
tf32_off (first untried candidate in allowlist order).[none, tf32_off]; none-none skipped
(flat resume); run tf32_off-none → still fails.xnack.xnack-none → if still fail and budget hit →
exhausted_candidates or policy_stop.If tf32_off-none passes (hypothetically):
Agent outcome: converged — Mitigation found — repro passes with a non-baseline cell.
Wrote /tmp/agent_out/smoke-fail/agent_report.md
Re-run the repro with mitigation `tf32_off` applied (see cell `tf32_off-none` probe.env or matrix).
Sample none-none/trial_0/result.json (abbreviated):
{
"verdict": "fail",
"exit_code": 1,
"cell_name": "none-none",
"failure_detectors_fired": ["tier1:exit_nonzero"],
"argv": ["python3", "-c", "import sys; sys.exit(1)"]
}
Command:
pip install 'amd-aorta[agent]'
export OPENAI_API_KEY=sk-...
PYTHONPATH=src aorta agent mitigate \
--output /tmp/agent_out \
--ticket smoke-llm \
--symptom "RCCL hang after checkpoint" \
--llm-backend litellm \
--llm-model gpt-4o-mini \
--mitigation none \
--mitigation nccl_launch_order_implicit \
--mitigation tf32_off \
-- \
./my_training_repro.sh
What happens under the hood:
{category, hypothesis, next_mitigations[], confidence, stop}.AgentPolicy.validate_step() drops any name not in the registry.The LLM may pick nccl_launch_order_implicit first because the symptom
mentions RCCL — unlike fake mode, which always takes the first untried name
in sorted allowlist order.
Re-run the same command with the same --output and --ticket:
PYTHONPATH=src aorta agent mitigate --output /tmp/agent_out --ticket smoke-fail -- ...
Under the hood:
wake() reads agent_log.jsonl and existing cell directories.run_recipe(..., resume_existing=True, layout="flat_resume") skips a
trial only when its own trial_<n>/result.json is complete (checked
per trial via aorta.probe.resume.is_trial_complete), so a cell reruns
the specific trials whose result is missing, incomplete, or corrupt —
not just when trial_0 is absent.--prompt-profile chooses the messages a real backend is sent. The loop, the
reply parser and AgentPolicy are the same under every profile.
| Profile | What the model is sent | Use it with |
|---|---|---|
default |
The agent’s own prompt: symptom, cell summaries, remaining candidates, already-tried list, and a gloss for each category | Any general-purpose model. Unchanged, and the default. |
rl-episode |
The prompt the probe policy is post-trained on in the RL episode environment (#525), byte for byte, with thinking disabled | Only a checkpoint trained on that prompt |
Why it exists. A post-trained checkpoint’s gain is tied to the words it
learned on. Measured through aorta agent mitigate --llm-backend vllm on seven
archived failure scenarios (six fixable), 8 runs each at the agent’s
temperature, counting first replies that name a mitigation which fixes the
failure:
| Model | default |
rl-episode |
|---|---|---|
| Qwen3-8B, base | 18 of 48 | 17 of 48 |
| Qwen3-8B, RL-trained (17 iterations) | 17 of 48 | 40 of 48 (5 of 6 scenarios) |
The RL-trained checkpoint was trained on all seven of these scenarios, so this table shows that the product path reproduces what the checkpoint learned under its training prompt. It is not evidence that the checkpoint generalises to failures it was not trained on.
Do not use it with a general model. It does not help the base model above,
and in 6 of the base model’s 48 runs the search ended early because it answered
with a category the loop does not accept. In an earlier sampled measurement,
Qwen3.8-27B named a fix on 49% of first replies under this prompt with thinking
off, against 61% under default with thinking allowed.
Serving. Serve the checkpoint with its Qwen3 reasoning parser, so the same
engine also answers default correctly. On TokenSpeed, also choose a sampling
backend that honours temperature (on AMD GPUs the default, greedy, ignores
it). rl-episode never sends JSON mode on either backend. Only default on
--llm-backend litellm does, so xgrammar matters only if the same engine
also serves that combination:
vllm serve /path/to/checkpoint --served-model-name aorta-probe \
--dtype bfloat16 --reasoning-parser qwen3
tokenspeed serve /path/to/checkpoint --served-model-name aorta-probe \
--dtype bfloat16 --reasoning-parser qwen3 \
--sampling-backend triton --grammar-backend xgrammar
Then point the agent at it through the chat provider settings:
export AORTA_CHAT_LLM_PROVIDER=vllm
export AORTA_CHAT_VLLM_BASE_URL=http://localhost:8000/v1
export AORTA_CHAT_VLLM_MODEL=aorta-probe
aorta agent mitigate --llm-backend vllm --prompt-profile rl-episode \
--output ./agent_results --ticket T1 -- ./my_repro.sh
What rl-episode changes besides the text:
chat_template_kwargs:
{"enable_thinking": false}), because the policy was trained without a
reasoning block. vLLM and TokenSpeed read this field; other servers may not.--symptom is not sent. The policy never saw one. What was already
tried is visible to it as the cells that ran, and tried names drop out of the
candidate list.[agent]-only LiteLLM path: training decoded
without a grammar.--llm-backend fake refuses it, since the fake proposer sends no prompt.verdict and detectors; the loop ignores
them.The rl-episode text lives in aorta/agent/prompt_profiles.py and is pinned
by digest in tests/agent/test_prompt_profiles.py. Editing it detaches every
checkpoint trained on it: replies still parse, but they get worse. Change it
only together with a checkpoint trained on the new text.
| Outcome | Meaning | Typical next step |
|---|---|---|
baseline_pass |
none-none passed |
No mitigations needed |
converged |
Some {mitigation}-none passed |
Ship that mitigation to customer / gate |
exhausted_candidates |
No mitigations left in allowlist/registry | Manual matrix or new sidecar mitigations |
agent_stop |
Proposer set stop (LLM or fake) |
Read agent_report.md hypothesis |
proposal_unresolved |
The proposer named mitigations, the candidate filter dropped all of them (unregistered, already tried, outside the allowlist, or the none baseline), and it did not ask to stop |
Check unresolved_mitigations in agent_log.jsonl against aorta mitigations list and --mitigation; do not read the hypothesis as the reason |
approval_required |
Mitigation needs ack (--require-approval) |
Operator approves, re-run |
walltime_exhausted |
--max-walltime-sec hit |
Re-run same ticket to resume |
policy_stop |
e.g. --max-iterations hit |
Increase budget or narrow allowlist |
aorta agent mitigate (CLI)
└── run_agent_loop() src/aorta/agent/loop.py
├── wake() replay agent_log.jsonl + cell verdicts
├── build_probe_recipe_from_dict()
├── run_recipe() same engine as aorta probe
│ └── SubprocessWorkload + 5-tier classifier
├── _read_cell_summaries() from trial_*/result.json
├── proposer.propose() fake OR LiteLLM
├── AgentPolicy.validate_step()
└── write_agent_report()
Every trial’s verdict comes from aorta.probe.classifier, not from the
agent:
custom_patternsSee classifier.md.
Proposed names must resolve via aorta.registry.get_mitigation(). Built-ins
include none, tf32_off, xnack, and many ROCm env-flag bundles in
src/aorta/registry/mitigations.py. Plugins register via the
aorta.mitigations entry-point group.
agent_log.jsonl)Append-only JSON lines, e.g.:
{"ts": "2026-06-04T12:00:00+00:00", "type": "session_start", "ticket": "smoke-hello", "llm_backend": "fake", ...}
{"ts": "...", "type": "llm_step", "category": "unknown", "hypothesis": "Baseline cell passed...", "stop": true, "stop_reason": "baseline_pass"}
{"ts": "...", "type": "search_stopped", "outcome": "baseline_pass", "stop_reason": "baseline_pass"}
llm_step and search_stopped carry an extra unresolved_mitigations key
only when the proposer named mitigations the candidate filter dropped:
{"ts": "...", "type": "llm_step", "next_mitigations": [], "stop": false, "stop_reason": null, "unresolved_mitigations": ["rccl_p2p_disable"]}
{"ts": "...", "type": "search_stopped", "outcome": "proposal_unresolved", "stop_reason": "proposal_unresolved", "unresolved_mitigations": ["rccl_p2p_disable"]}
The key is absent, not empty, when nothing was dropped — a run with no
rejections writes exactly the log it wrote before the key existed. It also
appears on a llm_step whose next_mitigations is non-empty, which is a
partial rejection: the search continued on the names that survived, and
this is the only record of the half that was discarded.
agent_report.md)One-page markdown: category, hypothesis, mitigation search table, evidence
chain (capture fields), recommended next action.
| Use fake (default) when… | Use litellm when… |
|---|---|
| CI, unit tests, offline dev | You want symptom-aware mitigation ordering |
| Reproducible demo | Large registry — LLM can prioritize likely fixes |
| No API keys / air-gapped | Richer hypotheses in agent_report.md |
Both backends share the same loop, policy, probe engine, and artifact layout.
aorta probe vs aorta agent mitigateaorta probe |
aorta agent mitigate |
|
|---|---|---|
| Matrix | You write full YAML axes | Grows axis iteration by iteration |
| Who picks next mitigation | You | Proposer (fake or LLM) |
| Verdict | Classifier | Classifier (unchanged) |
| argv | Opaque, fixed | Opaque, fixed |
| Resume | Per ticket dir | Same + agent_log.jsonl |
| LLM | Never | Optional at propose step only |
For a known matrix (regression gate), use aorta probe. For exploratory
“find a mitigation that makes this pass,” use aorta agent mitigate.
Confusing agent_stop message — upgrade to latest branch; baseline pass
should report baseline_pass with a clear success line.
Search does nothing after first run — ticket dir already has state; use a
fresh --ticket or inspect agent_log.jsonl.
ImportError: LiteLLM / does not provide the extra 'agent' — your venv
has an old aorta wheel (e.g. from PyPI) without the [agent] extra.
PYTHONPATH=src loads new agent code, but litellm was never installed.
From the aorta repo root on branch feature/aorta-probe-agent:
pip install -e '.[agent]'
# or, minimal fix:
pip install litellm
Then retry --llm-backend litellm.
All mitigations fail — expected for hard repros; outcome
exhausted_candidates; inspect failure_detectors_fired in
agent_report.md and consider manual probe matrix or new sidecar mitigations.