Bekzod Erkinov
Listen to Article
Loading...Table of contents · 8 sections
Claude Code Headless: Automate Dev Chores with claude -p
In claude -p, --allowedTools is an auto-approve list, not a restriction. I allowed only Write under bypassPermissions and a shell command still ran in 3 of 3 runs on Haiku and 3 of 3 on Sonnet.
The payoff: a permission matrix you can re-run with one script, a nightly diff-review script built on --tools, a stdout-versus-Write test for generating files, and the exit-code, timeout, memory and MCP-startup numbers I hit along the way.
Run on Claude Code 2.1.292, Windows (Git Bash), 2026-10-07, mostly --model haiku. Costs are the CLI's total_cost_usd, which the docs call a client-side estimate, not your bill.
Decision rule for unattended jobs
- Read-only job (review, summarize):
--tools "Read,Grep,Glob"+--permission-mode dontAsk+--strict-mcp-config --mcp-config '{"mcpServers":{}}'.- Job that must write:
--permission-mode acceptEdits, ordontAsk+ an exact allow-list such as--allowedTools "Bash(npm test)".bypassPermissions: only inside a throwaway container.
The basic call
"Headless" is just -p/--print: run one prompt, print the result, exit. The official headless page shows the shape:
claude -p "Find and fix the bug in auth.py" --allowedTools "Read,Edit,Bash"
cat build-error.txt | claude -p 'explain the root cause' > output.txt
Piped stdin is capped at 10MB per the docs. That's enough for a diff, not a log archive; write big inputs to a file and name the path in the prompt.
Also from the docs: without --bare, -p shows no workspace trust dialog, runs the hooks in the folder's .claude/settings.json and connects the servers in its .mcp.json. Only run it in directories you'd run code from.
claude -p flags: allowedTools vs tools vs disallowedTools
The flags people mix up are --permission-mode, --allowedTools, --disallowedTools and --tools. I asked the model to run echo hello > from_X.txt with Bash under each combination, 3 runs per row, and logged the stream with --output-format stream-json --verbose --forward-subagent-text. The last flag marks subagent messages with parent_tool_use_id, so the script can split Bash calls into main thread and subagent and count successful and errored results. A file appearing proves nothing alone, since the model routes around a blocked tool. This is the whole script (matrix.sh, needs claude and jq):
#!/usr/bin/env bash
# usage: matrix.sh <outdir> [row-ids...] (needs claude and jq on PATH)
out=$(realpath -m "${1:-matrix}"); mkdir -p "$out"; shift
echo '{"mcpServers":{}}' > "$out/empty.json"
NOMCP="--strict-mcp-config --mcp-config $out/empty.json"
P='Use the Bash tool to run: echo hello > from_$ROW.txt'
rows=(
"A|haiku|--permission-mode bypassPermissions --allowedTools Write"
"As|sonnet|--permission-mode bypassPermissions --allowedTools Write"
"B|haiku|--permission-mode bypassPermissions --tools Write"
"B2|haiku|--permission-mode bypassPermissions --tools Write $NOMCP"
"C|haiku|--permission-mode bypassPermissions --disallowedTools Bash"
"C2|haiku|--permission-mode bypassPermissions --disallowedTools Bash $NOMCP"
"D|haiku|--permission-mode bypassPermissions --disallowedTools Bash(echo*)"
"E|haiku|--permission-mode dontAsk --allowedTools Write"
"H|haiku|--permission-mode dontAsk --allowedTools Bash(echo*)"
"F|haiku|--permission-mode manual"
"G|haiku|"
)
for r in "${rows[@]}"; do
IFS='|' read -r id model flags <<< "$r"
[ $# -gt 0 ] && [[ " $* " != *" $id "* ]] && continue
for n in 1 2 3; do
d="$out/$id-$n"; mkdir -p "$d"
(cd "$d" && claude -p "${P/\$ROW/$id}" --model $model --no-session-persistence \
--output-format stream-json --verbose --forward-subagent-text $flags > s.jsonl 2> err.txt)
# Bash tool_use ids split by who made them: main thread (parent null) vs subagent (parent set)
cnt=$(jq -s '
[.[]|select(.type=="assistant")|.parent_tool_use_id as $p|.message.content[]?|select(.type=="tool_use" and .name=="Bash")|{id,sub:($p!=null)}] as $b
| [.[]|select(.type=="user")|.message.content[]?|select(.type=="tool_result")|{t:.tool_use_id,err:(.is_error//false)}] as $r
| def res(s): [$b[]|select(.sub==s)|.id as $i|$r[]|select(.t==$i)];
"main_ok=\(res(false)|map(select(.err|not))|length) main_err=\(res(false)|map(select(.err))|length) sub_ok=\(res(true)|map(select(.err|not))|length) sub_err=\(res(true)|map(select(.err))|length)"' "$d/s.jsonl" -r)
echo "$id run$n mode=$(jq -rs '[.[]|select(.type=="system" and .subtype=="init")|.permissionMode][0]' "$d/s.jsonl") $cnt file=$([ -f "$d/from_$id.txt" ] && echo yes || echo no) denials=$(jq -s '[.[]|select(.type=="result")|.permission_denials|length]|add' "$d/s.jsonl") main=$(jq -rs '[.[]|select(.type=="assistant" and .parent_tool_use_id==null)|.message.content[]?|select(.type=="tool_use")|.name|sub("^mcp__.*";"mcp")]|join(",")' "$d/s.jsonl") sub=$(jq -rs '[.[]|select(.type=="assistant" and .parent_tool_use_id!=null)|.message.content[]?|select(.type=="tool_use")|.name|sub("^mcp__.*";"mcp")]|join(",")' "$d/s.jsonl")"
done
done
Each row is 3 runs of one prompt, so the table checks the documented rules below; it is not a benchmark. Four lines of the script's output carry the key results (run 1 of rows A, B2, D and E; the other runs matched within each row except where the table says otherwise):
A run1 mode=bypassPermissions main_ok=1 main_err=0 file=yes denials=0 main=Bash
B2 run1 mode=bypassPermissions main_ok=0 main_err=1 file=yes denials=0 main=Bash,Write
D run1 mode=bypassPermissions main_ok=0 main_err=1 file=no denials=1 main=Bash
E run1 mode=dontAsk main_ok=0 main_err=1 file=yes denials=1 main=Bash,Write
(Trimmed from the script's sub_ok/sub_err/sub= columns, which were all zero or empty on these rows.)
| Row | Flags | Main-thread Bash succeeded | What the stream shows |
|---|---|---|---|
| A | bypassPermissions + --allowedTools Write |
3 of 3 (Sonnet As: 3 of 3) |
allow-list restricted nothing |
| B | bypassPermissions + --tools Write |
0 of 3 | Bash not offered; with MCP loaded the model called MCP tools in runs 1 and 2 |
| B2 | row B + empty strict MCP | 0 of 3 | the model tried Bash, got an error, then used Write |
| C | bypassPermissions + --disallowedTools Bash |
0 of 3 | Bash removed; the main thread called MCP tools, Skill and Agent |
| C2 | row C + empty strict MCP | 0 of 3 | main thread called only Agent |
| D | bypassPermissions + --disallowedTools "Bash(echo*)" |
0 of 3 | Bash attempted, errored, listed in permission_denials |
| E | dontAsk + --allowedTools Write |
0 of 3 | Bash denied every time; Write fallback created the file 3 of 3 |
| H | dontAsk + --allowedTools "Bash(echo*)" |
0 of 3 | denied; see the redirect test below |
| F | --permission-mode manual |
0 of 3 | denied (no prompt possible under -p) |
| G | no flags | 0 of 3 | same as F: permissionMode: default; manual is an alias |
What the output answers:
- Does
--disallowedTools Bashbind subagents? In the six C and C2 runs the subagents made noBashcall at all. They searched withToolSearchandSkill, and in 3 of 6 runs a subagent usedWriteto create the file. That fits Bash being unavailable to them, but a subagent that never tries Bash doesn't prove it couldn't, and none of my rows tests the mechanism. What the rows do show: a blocked Bash gets worked around with another tool. dontAskwith an allow-list works, when the rule matches. Row H was denied forecho hello > from_H.txt. My first follow-up changed two things at once, so I redid it with one change at a time (haiku,dontAsk, empty strict MCP, 3 runs each, scored byscore.pyon the stream):
rule command denied ran
Bash(echo *) echo hello > from_x.txt 3 of 3 0 of 3
Bash(echo*) echo hello > from_x.txt 3 of 3 0 of 3
Bash(echo*) echo hello 0 of 3 3 of 3
Same rule, same prompt, and only the > file part decides the outcome; the space in the pattern doesn't matter. So echo* and echo * both match a plain echo, and neither allow-rule covers an echo with a redirect under dontAsk. Row D's deny rule Bash(echo*) did match that redirect command, so deny rules appear to match more broadly than allow rules. That is what I observed, not something I found documented. If your job redirects output, allow the exact full command.
- With MCP loaded,
--toolsisn't the whole story. In B the main thread called MCP tools.--toolsdoesn't affect MCP tools; deny them with--disallowedTools "mcp__*"or run strict.
Row A is the one that bites. It isn't a bug; the permission-modes page says: "Deny rules block in every mode, including bypassPermissions... Allow rules have no effect in bypassPermissions." It also says bypassPermissions "offers no protection against prompt injection or unintended actions."
The rules:
--allowedToolsis an auto-approve list, not a restriction. Per the CLI reference: "To restrict which tools are available, use--toolsinstead."--disallowedToolsworks in every mode. A bare name likeBashremoves the tool; a scoped rule likeBash(rm *)leaves it and denies matching calls.dontAskdenies anything that would prompt. It still runs file reads in your working directories, the read-only command set, and anything your--allowedToolsentries cover. SodontAskgives you "never hangs";--toolsgives you "can't do X".- For unattended runs, combine them:
--permission-mode dontAskplus--tools(or an exact allow-list). The docs' CI example isclaude -p "run the test suite" --permission-mode dontAsk --allowedTools "Bash(npm test)" "Read". --permission-prompts none(v2.1.259+):claude --helpsays "nobody: anything that would prompt is denied automatically; the permission mode still decides everything else". The default ishost, the SDK host or--permission-prompt-tool.- Reserve
bypassPermissionsfor a throwaway container.claude --helpsays of--dangerously-skip-permissions: "Bypass all permission checks. Recommended only for sandboxes with no internet access."
The starting mode without flags: the permission-modes page, section Which mode a session starts in, gives claude -p the mode default in sessions that fetch feature flags. In sessions that don't (a third-party provider, telemetry off) it is auto on v2.1.285 or later and default before. My no-flags row reported default, so check permissionMode in system/init on your setup.
Denials don't fail the run: the denied runs above exited 0. Check permission_denials in the JSON output.
JSON output, stream-json and exit codes
--output-format json returns one object. These fields were present in my runs: result, session_id, total_cost_usd, usage, num_turns, duration_ms, is_error, subtype, terminal_reason, permission_denials. Add --json-schema and the validated data lands in structured_output (docs). Parse with jq:
out=$(claude -p "Summarize this diff" --output-format json < diff.patch)
echo "$out" | jq -r '.result'
echo "$out" | jq '{cost: .total_cost_usd, denied: (.permission_denials | length)}'
--output-format stream-json needs --verbose and emits one JSON object per line: hook events, system/init, assistant messages, a rate_limit_event, then a final result. The init event lists mcp_servers, tools and permissionMode, a handy place to assert what the run could touch.
Exit codes, as observed:
| Situation | Exit | Where the error is |
|---|---|---|
| Normal run | 0 | result |
| Unknown flag | 1 | stderr: error: unknown option '--bogus-flag' |
| Unknown model | 1 | stdout JSON, see below |
--max-turns 1 on a task needing a tool |
1 | stdout JSON, see below |
| Tool denied by permissions | 0 | permission_denials array |
timeout 3 claude -p ... |
124 | killed from outside |
Raw output for the two JSON rows (--model not-a-model, and --max-turns 1 with a prompt that needs Bash), picked from the --output-format json object:
unknown model rc=1 {'subtype': 'success', 'is_error': True, 'terminal_reason': 'api_error', 'result': "There's an issue with the selected model (not-a-model). It may not exist or you may not have access to it. ..."}
max-turns 1 rc=1 {'subtype': 'error_max_turns', 'is_error': True, 'terminal_reason': 'max_turns', 'result': None, 'num_turns': 2}
An unknown model reports subtype: "success" next to is_error: true. Gate on is_error and the exit code, not subtype. A run that couldn't do its job because of a denial still exits 0, so in CI also fail on a non-empty permission_denials.
Generating files: print to stdout or let the model Write?
A chore that produces a document has two options: have claude -p print it and redirect stdout, or give it Write and a path. I ran both once on Haiku with the same brief (a 4,000-word PostgreSQL indexing guide), --output-format json, empty strict MCP:
stdout (--tools "") |
Write (--tools Write, dontAsk) |
|
|---|---|---|
| Exit code | 0 | 0 |
| Wall time | 70 s | 65 s |
| Turns | 1 | 2 |
| Output tokens | 5,209 | 5,163 |
total_cost_usd |
$0.052 | $0.065 |
| Words delivered | 3,198 in result |
2,776 in guide.md, result was just DONE |
stop_reason |
end_turn |
end_turn |
Nothing was truncated at this size, and both fell short of 4,000 words (single runs, so don't read anything into 3,198 versus 2,776). The difference is in the stdout copy: it opened with I'll write a comprehensive practical guide to PostgreSQL indexing covering all your requested topics. and a --- line before the title. Redirect that to a .md file and the chatter is now your first line. The Write file began at # A Practical Guide.... The price is one extra turn and about 25% more cost here.
Decision rule: if the output is a file that other tools consume, give the model Write and a fixed path (and check the file exists afterwards), because stdout carries preamble you would have to strip. If it's a short answer a script reads, such as a verdict, a label or a commit message, use stdout with --json-schema so the shape is validated. I didn't test outputs near the model's output limit, so I can't tell you where truncation starts.
Claude Code headless timeouts: parallel runs were slower
A trivial say ok on Haiku took 11-18 s solo against about 6 s of API time (duration_api_ms); the rest is startup (settings, hooks, plugins, MCP servers). Then I varied concurrency with conc.sh (default config, N processes launched together):
N=1: 13.7 s
N=2: 14.2 15.1 s
N=4: 15.9 21.5 21.8 23.2 s
N=6: 24.5 26.3 27.3 27.5 28.5 29.7 s
These are single runs per N on one laptop, so treat the trend as indicative. Per-process wall time climbs with N: the median is about 14 s at N=2, 21 s at N=4 and 27 s at N=6. Repeat N=6 runs landed between 20 and 31 s. An earlier N=6 set measured 37.7 to 54.6 s, which contradicts these; I never found out why (I didn't record what else was running), so I'm treating it as an unexplained outlier and not using it. Budget for swings of 2x, not several seconds.
For memory I sampled two ways during N=6 runs (machine total 32,601 MB). System-wide Win32_OperatingSystem.FreePhysicalMemory, once a second, fell from 19,313 MB to a minimum of 17,041 MB: a 2.3 GB drop with other apps open. Then procmem.ps1 polled every 500 ms and summed WorkingSet64 of every process that didn't exist before the run:
22.5 23.4 23.8 24.6 25.1 31.0
peak new-process working set: 3,833 MB total, 1,675 MB in claude.exe
Per run that is about 280 MB in claude.exe and about 640 MB counting every child process (hooks, MCP servers). Working set counts shared pages in each process, so it overstates; the free-memory drop (about 380 MB per run) is a system-wide sample. Budget between the two. I didn't sample CPU, so I can't say whether the slowdown is CPU, disk or API-side. Practical consequences:
- Set an outer timeout (
timeout 600 claude -p ..., exit 124 on expiry) sized to your slow case, not the median. - Add
--max-turnsand--max-budget-usdas inner brakes. - Don't launch N runs with
&and expect N-times throughput. Cap concurrency at 2-3. - Per the docs, a dev server the model starts is killed about 5 s after the result; a background subagent keeps
-popen until it finishes or a 10-minute idle ceiling hits.
--strict-mcp-config: skip the cold start
My user config has plugin and claude.ai MCP servers. A default run's init event listed 25 of them (connected, pending, needs-auth, failed) and 171 tools. With --strict-mcp-config --mcp-config '{"mcpServers":{}}' it listed no servers and 32 tools. The CLI reference says the flag means "Only use MCP servers from --mcp-config, ignoring all other MCP configurations."
Same say ok prompt, Haiku, three runs per config, interleaved back to back, same working directory:
| Config | Wall time (3 runs) | Median | Cost (3 runs) |
|---|---|---|---|
| Default | 12.3, 9.0, 11.9 s | 11.9 s | $0.042, $0.005, $0.005 |
--strict-mcp-config (empty) |
6.1, 12.7, 6.8 s | 6.8 s | $0.033, $0.007, $0.004 |
--strict-mcp-config + --safe-mode |
3.5, 3.4, 3.7 s | 3.5 s | $0.009, $0.002, $0.003 |
Run 1 of each config paid for cache creation, so compare runs 2 and 3: all three configs are fractions of a cent. The wall-time gap is the more reliable effect, though strict run 2 (12.7 s) shows it's noisy. --safe-mode also disables CLAUDE.md, skills, hooks and plugins, so it's a troubleshooting flag, not a production one.
Two things the docs add that bite in scripts:
- With
--mcp-configand-p, Claude waits for still-pending servers before the first turn, up toMCP_TIMEOUT(30 s by default, v2.1.221+). A slow server adds that wait to every run. - Config entries that fail validation are skipped, and the run still exits
0. I passed{"bad":{"url":"http://localhost:9/mcp"}}andsystem/initreported"mcp_servers":[]and"mcp_server_errors":[{"name":"bad","type":"url_missing_type","message":"Skipped — MCP server \"bad\" has a \"url\" but no \"type\"; add \"type\": \"http\" (or \"sse\" / \"ws\") to this entry"}]. If a job needs an MCP server, fail it whenmcp_server_errorsis non-empty (v2.1.219+).
A nightly review that has something to review
My first version reviewed git diff --staged. At 02:00 nothing is staged, so a cron job running it reviews an empty diff. A scheduled job needs input that exists without you: here, whatever landed on origin/main in the last 24 hours.
I built a bare origin, a commit dated three days ago, and a second commit from now (get_page made 1-based, plus a new last_page). Then I cloned a "server" copy that had only seen the old commit and ran this against it:
#!/usr/bin/env bash
# usage: nightly.sh /path/to/repo -- review what landed on origin/main in the last 24 h
set -u
cd "$1" || exit 1
git fetch -q origin || { echo "git fetch failed"; exit 1; }
base=$(git rev-list -1 --before='24 hours ago' origin/main)
[ -n "$base" ] || base=$(git hash-object -t tree /dev/null) # young repo: review everything
diff=$(git diff "$base" origin/main) || { echo "git diff failed"; exit 1; }
[ -n "$diff" ] || { echo "$(date -Is) nothing landed in 24h"; exit 0; }
out=$(printf '%s\n' "$diff" | timeout 600 claude -p "Review this diff for bugs. Report file:line and one sentence each." \
--model haiku \
--tools "Read,Grep,Glob" \
--permission-mode dontAsk \
--strict-mcp-config --mcp-config '{"mcpServers":{}}' \
--max-turns 15 --max-budget-usd 1.00 \
--no-session-persistence \
--output-format json)
rc=$? # status of the pipeline's last command (timeout/claude); printf's is not seen
[ $rc -ne 0 ] && { echo "claude failed rc=$rc"; exit $rc; }
echo "$out" | jq -e '.is_error == false and (.permission_denials | length) == 0' >/dev/null \
|| { echo "$out" | jq -r '.result'; exit 1; }
echo "$(date -Is) review of $base..origin/main"
echo "$out" | jq -r '.result'
rc=$? holds the exit status of the last command in the pipeline (timeout ... claude), so a failure of printf wouldn't show up there; that's why git diff is checked separately, where its failure is reported instead of silently reviewing an empty string.
Why is --tools "Read,Grep,Glob" enough, given the MCP finding above? The empty strict config removes MCP, and --tools removes everything else that isn't listed, Agent included. The init event of a run with exactly these flags showed "tools":["Glob","Grep","Read"] and "mcp_servers":[]: nothing left that can run a command or spawn a subagent.
There is no cron daemon in Git Bash on Windows, so I simulated its environment: env -i with a bare PATH, the cron command line from below (>> review.log 2>&1), from /. An earlier attempt with git missing from that PATH failed the way cron jobs do:
nightly.sh: line 5: git: command not found
git fetch failed
Cron's default PATH is minimal; set it in the crontab. With git, claude and jq on the path, the script exited 0 in about 26 s and review.log ended with:
2026-10-07T03:31:19+05:00 review of 30382599eaa2728a72d1ed8a00d206ba6c0893da..origin/main
**pager.py:8** — `last_page()` uses `page_count()` which only counts complete pages via integer division, so it returns an incomplete page when items aren't evenly divisible by `page_size` (e.g., 10 items with page_size 3 has 4 pages but `page_count()` returns 3).
That's the planted bug: last_page calls the existing floor-division page_count. An earlier run also flagged the pre-existing page_count line, so treat line numbers as pointers, not citations. The log also held two SessionEnd hook ... failed lines (node: command not found under the stripped PATH). They come from my personal user hooks, which -p runs without --bare: a hook in your environment can write into a "clean" log.
The crontab line:
chmod +x nightly.sh
# crontab -e: 02:00 daily, append output to review.log
PATH=/usr/local/bin:/usr/bin:/bin
0 2 * * * /path/to/nightly.sh /path/to/repo >> /path/to/review.log 2>&1
What to do next
- Run
matrix.shagainst your own model and config. It takes about ten minutes and shows which routes to other tools are open in your setup (MCP servers and subagents were the surprise in mine). - Gate CI on the output: fail the job when
is_erroris true,permission_denialsis non-empty, or (if the job needs MCP)mcp_server_errorsis non-empty. - Pin a timeout and budget:
timeout 600around the call,timeout-minuteson the CI job,--max-turnsplus--max-budget-usdinside, and RAM headroom of roughly 300 to 650 MB per concurrent run (the two measurements above). - Prefer
--barefor scripts once you have anANTHROPIC_API_KEY. The docs call it the recommended mode for scripted and SDK calls; it skips hooks, plugins, MCP discovery and CLAUDE.md.
More on the tooling around chores like these: 5 Essential Tools for Every Full-Stack Developer and Real-World Examples of AI Integration in Web Development. The full list is in the AI Tutorials category.
Not tested: --bare (no API key here), --restricted (v2.1.248+; claude --help says it removes Bash, PowerShell, REPL and other code-running tools, plus WebFetch), SIGTERM exit 143 (docs-only), cron mail behaviour, a real cron daemon, GitHub Actions (see the docs' GitHub Actions page), and generation near the output-token limit.
Keep reading
Decentralized Finance Protocol Comparison: Uniswap V3 vs SushiSwap vs Curve Finance - Performance, Security, and Use Cases
16 min · 718 views
AI WorkflowsMulti-Tenant SaaS Architecture with Laravel and React: Battle-Tested Patterns from Production
33 min · 152 views
AI WorkflowsPostgreSQL Hidden Features for Better Performance: Battle-Tested Optimizations from Production
24 min · 133 views
Bekzod Erkinov
AuthorFounder of NextGenBeing. Software engineer working with Laravel, Python, and cloud infrastructure. Writes about patterns that actually hold up in production. Based in Tashkent, Uzbekistan.
Get the AI-Assisted Developer's Field Guide
The workflow, prompts, and tools I use to ship faster with AI — free when you subscribe. Plus new deep-dives in your inbox. No spam, unsubscribe anytime.
Comments (0)
Please log in to leave a comment.
Log InRelated Articles
Building a Complete E-commerce Website with Laravel: What We Learned Scaling to 100k Orders
Apr 30, 2026
PostgreSQL Hidden Features for Better Performance: Battle-Tested Optimizations from Production
May 23, 2026
Decentralized Finance Protocol Comparison: Uniswap V3 vs SushiSwap vs Curve Finance - Performance, Security, and Use Cases
Feb 28, 2026