Mentu

CLI agents

CLI agents

Two backends run a step by shelling out to an agent CLI already installed on your machine.

Backend Requires Stream System context
claude a claude executable on PATH claude_json passed natively
codex a codex executable on PATH codex_json folded into the prompt

Both report as local, and neither needs the runner to hold a credential: they use whatever the CLI is already signed in with.

mentu-recipes adapters --explain codex
codex
  execution: agent-cli
  stream: codex_json
  completion: provider_complete_event
  local: true
  network: false
  credential: false
  tools: true
  structured completion: true

The network: false line describes the runner, not the CLI. The runner opens no socket. Whether the CLI it launches reaches the network is the CLI's business.

What the runner invokes

Both CLIs are launched with their permission prompts disabled, so the agent can edit files and run commands in the workspace without stopping to ask. Treat a recipe that uses these backends the way you would treat a script with the same access.

The Claude CLI is run as claude -p <prompt> --dangerously-skip-permissions --verbose --output-format stream-json, with:

  • --model from the step, command line, or recipe model
  • --effort from reasoning, and --thinking from thinking
  • --append-system-prompt pointing at a file holding the step's system context
  • --allowedTools from allowed_tools
  • --disallowedTools from disallowed_tools, always joined with TodoWrite, TaskCreate, TaskUpdate, TaskList, and TaskGet

A prompt over 7000 bytes is written to .mentu/tmp/ inside the workspace and the CLI is told to read it from there.

The Codex CLI is run as codex exec --json --color never --dangerously-bypass-approvals-and-sandbox -s danger-full-access -C <workspace> --skip-git-repo-check <prompt>, with:

  • --model from the step, command line, or recipe model
  • model_reasoning_effort from reasoning, with max mapped to xhigh
  • --output-last-message pointing at a file under .mentu/tmp/, used as the step output when the JSON stream carried no message text

System context for Codex is prepended to the prompt under a Context: heading, which is what folded_into_prompt means in the adapters table. thinking is not passed to Codex; doctor reports unsupported_thinking if a Codex step sets it.

A step against a CLI agent

{
  "label": "edit",
  "backend": "claude",
  "prompt_file": "PROMPT-tidy.md",
  "completion_keyword": "TIDY_COMPLETE",
  "expected_changes": ["CHANGELOG.md"],
  "allowed_tools": ["Read", "Edit"],
  "timeout": 300
}

An agent CLI edits files, so this is where expected_changes earns its place. When the step completes, files matching the declaration are staged and committed with the message chore: mentu-recipes step <label> (<run-id>). Anything else the agent touched is written to a patch under the run's quarantine/ directory and left out of the commit. If the agent wrote only outside the declaration, the step fails. This is what you want when a model wanders outside the task.

End the prompt with an instruction to print the completion keyword. Without one, doctor reports missing_completion_signal on any non-shell step that has neither completion_keyword nor verify: a step that never says it finished cannot be checked.

Tool allow-lists

allowed_tools and disallowed_tools reach the Claude CLI as --allowedTools and --disallowedTools. The Codex adapter does not pass them: codex exec has no equivalent flags, and Codex runs with its sandbox bypassed as described above. Both backends report supports_tool_allow_list and supports_tool_deny_list as true in mentu-recipes adapters --json, so doctor does not warn about tool lists on a Codex step. It does warn, with unsupported_allowed_tools or unsupported_disallowed_tools, when they are set on an HTTP or shell backend.

Environment isolation

Each CLI gets an allow-listed environment, not your whole shell. The runner starts from your process environment merged with the recipe and step env blocks, then keeps only PATH, HOME, SHELL, USER, LOGNAME, TMPDIR, TMP, TEMP, TERM, LANG, LC_ALL and other LC_* variables, SSL_CERT_FILE, SSL_CERT_DIR, the proxy variables, and XDG_*, plus exactly one set of credentials per backend: ANTHROPIC_API_KEY for the Claude CLI, OPENAI_API_KEY and CODEX_HOME for the Codex CLI. An OpenAI key is not handed to Claude, and an Anthropic key is not handed to Codex. Nothing is read from the Keychain on the CLIs' behalf; a vault key reaches them only through an env reference, see Credentials.

Output handling

Both adapters parse the CLI's JSON stream and store the assistant's message text as the step output in the run record, with token counts when the CLI reports them. Provider log lines are filtered out of the stream you watch and out of the stored stderr. Lines that look actionable, such as missing auth, rate limits, quota or billing errors, and unsupported options, are kept and shown as Error: lines.

Choosing the agent for a custom cli provider

A named provider with "api": "cli" drives an agent CLI installed on the Mac. Which one is chosen by the agent field: claude when the field is absent, codex, or pi. There is no per-vendor API value; the provider surface stays the same for every agent.

"providers": {
  "local-agent": {
    "api": "cli",
    "agent": "pi",
    "base_url": "http://127.0.0.1:8080/v1",
    "api_key_env": "LOCAL_AGENT_KEY",
    "model": "your-exact-model-id"
  }
}

pi runs the Pi coding agent in a private profile, reads the prompt from a file on stdin, needs Pi 0.84.1 and Node 22.19 or newer, is never selected automatically, and reports Pi's own token usage into input_tokens and output_tokens in the run record. Full details, including tool names and the max_tokens_field, context_window and skills fields: docs/pi-backend.md.

© 2026 Mentu.