CLI agents
CLI agents
Two backends run a step by shelling out to an agent CLI already installed on your machine.
| Backend | Requires | Stream | System context |
|---|---|---|---|
claude |
a claude executable on PATH |
claude_json |
passed natively |
codex |
a codex executable on PATH |
codex_json |
folded into the prompt |
Both report as local, and neither needs the runner to hold a credential: they use whatever the CLI is already signed in with.
mentu-recipes adapters --explain codexcodex
execution: agent-cli
stream: codex_json
completion: provider_complete_event
local: true
network: false
credential: false
tools: true
structured completion: trueThe network: false line describes the runner, not the CLI. The runner opens no
socket. Whether the CLI it launches reaches the network is the CLI's business.
What the runner invokes
Both CLIs are launched with their permission prompts disabled, so the agent can edit files and run commands in the workspace without stopping to ask. Treat a recipe that uses these backends the way you would treat a script with the same access.
The Claude CLI is run as claude -p <prompt> --dangerously-skip-permissions --verbose --output-format stream-json, with:
--modelfrom the step, command line, or recipemodel--effortfromreasoning, and--thinkingfromthinking--append-system-promptpointing at a file holding the step's system context--allowedToolsfromallowed_tools--disallowedToolsfromdisallowed_tools, always joined withTodoWrite,TaskCreate,TaskUpdate,TaskList, andTaskGet
A prompt over 7000 bytes is written to .mentu/tmp/ inside the workspace and the
CLI is told to read it from there.
The Codex CLI is run as codex exec --json --color never --dangerously-bypass-approvals-and-sandbox -s danger-full-access -C <workspace> --skip-git-repo-check <prompt>, with:
--modelfrom the step, command line, or recipemodelmodel_reasoning_effortfromreasoning, withmaxmapped toxhigh--output-last-messagepointing at a file under.mentu/tmp/, used as the step output when the JSON stream carried no message text
System context for Codex is prepended to the prompt under a Context: heading,
which is what folded_into_prompt means in the adapters table. thinking is
not passed to Codex; doctor reports unsupported_thinking if a Codex step
sets it.
A step against a CLI agent
{
"label": "edit",
"backend": "claude",
"prompt_file": "PROMPT-tidy.md",
"completion_keyword": "TIDY_COMPLETE",
"expected_changes": ["CHANGELOG.md"],
"allowed_tools": ["Read", "Edit"],
"timeout": 300
}An agent CLI edits files, so this is where expected_changes earns its place.
When the step completes, files matching the declaration are staged and committed
with the message chore: mentu-recipes step <label> (<run-id>). Anything else
the agent touched is written to a patch under the run's quarantine/ directory
and left out of the commit. If the agent wrote only outside the declaration, the
step fails. This is what you want when a model wanders outside the task.
End the prompt with an instruction to print the completion keyword. Without one,
doctor reports missing_completion_signal on any non-shell step that has
neither completion_keyword nor verify: a step that never says it finished
cannot be checked.
Tool allow-lists
allowed_tools and disallowed_tools reach the Claude CLI as --allowedTools
and --disallowedTools. The Codex adapter does not pass them: codex exec has
no equivalent flags, and Codex runs with its sandbox bypassed as described
above. Both backends report supports_tool_allow_list and
supports_tool_deny_list as true in mentu-recipes adapters --json, so
doctor does not warn about tool lists on a Codex step. It does warn, with
unsupported_allowed_tools or unsupported_disallowed_tools, when they are set
on an HTTP or shell backend.
Environment isolation
Each CLI gets an allow-listed environment, not your whole shell. The runner
starts from your process environment merged with the recipe and step env
blocks, then keeps only PATH, HOME, SHELL, USER, LOGNAME, TMPDIR,
TMP, TEMP, TERM, LANG, LC_ALL and other LC_* variables,
SSL_CERT_FILE, SSL_CERT_DIR, the proxy variables, and XDG_*, plus exactly
one set of credentials per backend: ANTHROPIC_API_KEY for the Claude CLI,
OPENAI_API_KEY and CODEX_HOME for the Codex CLI. An OpenAI key is not
handed to Claude, and an Anthropic key is not handed to Codex. Nothing is read
from the Keychain on the CLIs' behalf; a vault key reaches them only through an
env reference, see Credentials.
Output handling
Both adapters parse the CLI's JSON stream and store the assistant's message text
as the step output in the run record, with token counts when the CLI reports
them. Provider log lines are filtered out of the stream you watch and out of the
stored stderr. Lines that look actionable, such as missing auth, rate limits,
quota or billing errors, and unsupported options, are kept and shown as
Error: lines.
Choosing the agent for a custom cli provider
A named provider with "api": "cli" drives an agent CLI installed on the Mac.
Which one is chosen by the agent field: claude when the field is absent,
codex, or pi. There is no per-vendor API value; the provider surface stays
the same for every agent.
"providers": {
"local-agent": {
"api": "cli",
"agent": "pi",
"base_url": "http://127.0.0.1:8080/v1",
"api_key_env": "LOCAL_AGENT_KEY",
"model": "your-exact-model-id"
}
}pi runs the Pi coding agent in a private profile, reads the prompt from a
file on stdin, needs Pi 0.84.1 and Node 22.19 or newer, is never selected
automatically, and reports Pi's own token usage into input_tokens and
output_tokens in the run record. Full details, including tool names and the
max_tokens_field, context_window and skills fields:
docs/pi-backend.md.