Claude Code Headless: Automate Dev Chores with claude -p - NextGenBeing
Back to discoveries

Claude Code Headless: Automate Dev Chores with claude -p

Run claude -p headless for nightly jobs and CI: which permission flags really restrict tools, exit codes, JSON output, timeouts, RAM use and MCP cold starts.

AI Workflows 12 min read
Bekzod Erkinov

Bekzod Erkinov

Oct 6, 2026 • 0 views
Claude Code Headless: Automate Dev Chores with claude -p
Photo by Albert Stoynov on Unsplash
Size:
Height:
📖 12 min read 📝 4,891 words 👁 Focus mode: ✨ Eye care:

Listen to Article

Loading...
0:00 / 0:00
0:00 0:00
Low High
0% 100%
⏸ Paused ▶️ Now playing... Ready to play ✓ Finished
Table of contents · 8 sections

Claude Code Headless: Automate Dev Chores with claude -p

In claude -p, --allowedTools is an auto-approve list, not a restriction. I allowed only Write under bypassPermissions and a shell command still ran in 3 of 3 runs on Haiku and 3 of 3 on Sonnet.

The payoff: a permission matrix you can re-run with one script, a nightly diff-review script built on --tools, a stdout-versus-Write test for generating files, and the exit-code, timeout, memory and MCP-startup numbers I hit along the way.

Run on Claude Code 2.1.292, Windows (Git Bash), 2026-10-07, mostly --model haiku. Costs are the CLI's total_cost_usd, which the docs call a client-side estimate, not your bill.

Decision rule for unattended jobs

  • Read-only job (review, summarize): --tools "Read,Grep,Glob" + --permission-mode dontAsk + --strict-mcp-config --mcp-config '{"mcpServers":{}}'.
  • Job that must write: --permission-mode acceptEdits, or dontAsk + an exact allow-list such as --allowedTools "Bash(npm test)".
  • bypassPermissions: only inside a throwaway container.

The basic call

"Headless" is just -p/--print: run one prompt, print the result, exit. The official headless page shows the shape:

claude -p "Find and fix the bug in auth.py" --allowedTools "Read,Edit,Bash"
cat build-error.txt | claude -p 'explain the root cause' > output.txt

Piped stdin is capped at 10MB per the docs. That's enough for a diff, not a log archive; write big inputs to a file and name the path in the prompt.

Also from the docs: without --bare, -p shows no workspace trust dialog, runs the hooks in the folder's .claude/settings.json and connects the servers in its .mcp.json. Only run it in directories you'd run code from.

claude -p flags: allowedTools vs tools vs disallowedTools

The flags people mix up are --permission-mode, --allowedTools, --disallowedTools and --tools. I asked the model to run echo hello > from_X.txt with Bash under each combination, 3 runs per row, and logged the stream with --output-format stream-json --verbose --forward-subagent-text. The last flag marks subagent messages with parent_tool_use_id, so the script can split Bash calls into main thread and subagent and count successful and errored results. A file appearing proves nothing alone, since the model routes around a blocked tool. This is the whole script (matrix.sh, needs claude and jq):

#!/usr/bin/env bash
# usage: matrix.sh <outdir> [row-ids...]   (needs claude and jq on PATH)
out=$(realpath -m "${1:-matrix}"); mkdir -p "$out"; shift
echo '{"mcpServers":{}}' > "$out/empty.json"
NOMCP="--strict-mcp-config --mcp-config $out/empty.json"
P='Use the Bash tool to run: echo hello > from_$ROW.txt'
rows=(
 "A|haiku|--permission-mode bypassPermissions --allowedTools Write"
 "As|sonnet|--permission-mode bypassPermissions --allowedTools Write"
 "B|haiku|--permission-mode bypassPermissions --tools Write"
 "B2|haiku|--permission-mode bypassPermissions --tools Write $NOMCP"
 "C|haiku|--permission-mode bypassPermissions --disallowedTools Bash"
 "C2|haiku|--permission-mode bypassPermissions --disallowedTools Bash $NOMCP"
 "D|haiku|--permission-mode bypassPermissions --disallowedTools Bash(echo*)"
 "E|haiku|--permission-mode dontAsk --allowedTools Write"
 "H|haiku|--permission-mode dontAsk --allowedTools Bash(echo*)"
 "F|haiku|--permission-mode manual"
 "G|haiku|"
)
for r in "${rows[@]}"; do
  IFS='|' read -r id model flags <<< "$r"
  [ $# -gt 0 ] && [[ " $* " != *" $id "* ]] && continue
  for n in 1 2 3; do
    d="$out/$id-$n"; mkdir -p "$d"
    (cd "$d" && claude -p "${P/\$ROW/$id}" --model $model --no-session-persistence \
       --output-format stream-json --verbose --forward-subagent-text $flags > s.jsonl 2> err.txt)
    # Bash tool_use ids split by who made them: main thread (parent null) vs subagent (parent set)
    cnt=$(jq -s '
      [.[]|select(.type=="assistant")|.parent_tool_use_id as $p|.message.content[]?|select(.type=="tool_use" and .name=="Bash")|{id,sub:($p!=null)}] as $b
      | [.[]|select(.type=="user")|.message.content[]?|select(.type=="tool_result")|{t:.tool_use_id,err:(.is_error//false)}] as $r
      | def res(s): [$b[]|select(.sub==s)|.id as $i|$r[]|select(.t==$i)];
        "main_ok=\(res(false)|map(select(.err|not))|length) main_err=\(res(false)|map(select(.err))|length) sub_ok=\(res(true)|map(select(.err|not))|length) sub_err=\(res(true)|map(select(.err))|length)"' "$d/s.jsonl" -r)
    echo "$id run$n mode=$(jq -rs '[.[]|select(.type=="system" and .subtype=="init")|.permissionMode][0]' "$d/s.jsonl") $cnt file=$([ -f "$d/from_$id.txt" ] && echo yes || echo no) denials=$(jq -s '[.[]|select(.type=="result")|.permission_denials|length]|add' "$d/s.jsonl") main=$(jq -rs '[.[]|select(.type=="assistant" and .parent_tool_use_id==null)|.message.content[]?|select(.type=="tool_use")|.name|sub("^mcp__.*";"mcp")]|join(",")' "$d/s.jsonl") sub=$(jq -rs '[.[]|select(.type=="assistant" and .parent_tool_use_id!=null)|.message.content[]?|select(.type=="tool_use")|.name|sub("^mcp__.*";"mcp")]|join(",")' "$d/s.jsonl")"
  done
done

Each row is 3 runs of one prompt, so the table checks the documented rules below; it is not a benchmark. Four lines of the script's output carry the key results (run 1 of rows A, B2, D and E; the other runs matched within each row except where the table says otherwise):

A  run1 mode=bypassPermissions main_ok=1 main_err=0 file=yes denials=0 main=Bash
B2 run1 mode=bypassPermissions main_ok=0 main_err=1 file=yes denials=0 main=Bash,Write
D  run1 mode=bypassPermissions main_ok=0 main_err=1 file=no  denials=1 main=Bash
E  run1 mode=dontAsk           main_ok=0 main_err=1 file=yes denials=1 main=Bash,Write

(Trimmed from the script's sub_ok/sub_err/sub= columns, which were all zero or empty on these rows.)

Row Flags Main-thread Bash succeeded What the stream shows
A bypassPermissions + --allowedTools Write 3 of 3 (Sonnet As: 3 of 3) allow-list restricted nothing
B bypassPermissions + --tools Write 0 of 3 Bash not offered; with MCP loaded the model called MCP tools in runs 1 and 2
B2 row B + empty strict MCP 0 of 3 the model tried Bash, got an error, then used Write
C bypassPermissions + --disallowedTools Bash 0 of 3 Bash removed; the main thread called MCP tools, Skill and Agent
C2 row C + empty strict MCP 0 of 3 main thread called only Agent
D bypassPermissions + --disallowedTools "Bash(echo*)" 0 of 3 Bash attempted, errored, listed in permission_denials
E dontAsk + --allowedTools Write 0 of 3 Bash denied every time; Write fallback created the file 3 of 3
H dontAsk + --allowedTools "Bash(echo*)" 0 of 3 denied; see the redirect test below
F --permission-mode manual 0 of 3 denied (no prompt possible under -p)
G no flags 0 of 3 same as F: permissionMode: default; manual is an alias

What the output answers:

  • Does --disallowedTools Bash bind subagents? In the six C and C2 runs the subagents made no Bash call at all. They searched with ToolSearch and Skill, and in 3 of 6 runs a subagent used Write to create the file. That fits Bash being unavailable to them, but a subagent that never tries Bash doesn't prove it couldn't, and none of my rows tests the mechanism. What the rows do show: a blocked Bash gets worked around with another tool.
  • dontAsk with an allow-list works, when the rule matches. Row H was denied for echo hello > from_H.txt. My first follow-up changed two things at once, so I redid it with one change at a time (haiku, dontAsk, empty strict MCP, 3 runs each, scored by score.py on the stream):
rule             command                    denied  ran
Bash(echo *)     echo hello > from_x.txt    3 of 3  0 of 3
Bash(echo*)      echo hello > from_x.txt    3 of 3  0 of 3
Bash(echo*)      echo hello                 0 of 3  3 of 3

Same rule, same prompt, and only the > file part decides the outcome; the space in the pattern doesn't matter. So echo* and echo * both match a plain echo, and neither allow-rule covers an echo with a redirect under dontAsk. Row D's deny rule Bash(echo*) did match that redirect command, so deny rules appear to match more broadly than allow rules. That is what I observed, not something I found documented. If your job redirects output, allow the exact full command.

  • With MCP loaded, --tools isn't the whole story. In B the main thread called MCP tools. --tools doesn't affect MCP tools; deny them with --disallowedTools "mcp__*" or run strict.

Row A is the one that bites. It isn't a bug; the permission-modes page says: "Deny rules block in every mode, including bypassPermissions... Allow rules have no effect in bypassPermissions." It also says bypassPermissions "offers no protection against prompt injection or unintended actions."

The rules:

  • --allowedTools is an auto-approve list, not a restriction. Per the CLI reference: "To restrict which tools are available, use --tools instead."
  • --disallowedTools works in every mode. A bare name like Bash removes the tool; a scoped rule like Bash(rm *) leaves it and denies matching calls.
  • dontAsk denies anything that would prompt. It still runs file reads in your working directories, the read-only command set, and anything your --allowedTools entries cover. So dontAsk gives you "never hangs"; --tools gives you "can't do X".
  • For unattended runs, combine them: --permission-mode dontAsk plus --tools (or an exact allow-list). The docs' CI example is claude -p "run the test suite" --permission-mode dontAsk --allowedTools "Bash(npm test)" "Read".
  • --permission-prompts none (v2.1.259+): claude --help says "nobody: anything that would prompt is denied automatically; the permission mode still decides everything else". The default is host, the SDK host or --permission-prompt-tool.
  • Reserve bypassPermissions for a throwaway container. claude --help says of --dangerously-skip-permissions: "Bypass all permission checks. Recommended only for sandboxes with no internet access."

The starting mode without flags: the permission-modes page, section Which mode a session starts in, gives claude -p the mode default in sessions that fetch feature flags. In sessions that don't (a third-party provider, telemetry off) it is auto on v2.1.285 or later and default before. My no-flags row reported default, so check permissionMode in system/init on your setup.

Denials don't fail the run: the denied runs above exited 0. Check permission_denials in the JSON output.

JSON output, stream-json and exit codes

--output-format json returns one object. These fields were present in my runs: result, session_id, total_cost_usd, usage, num_turns, duration_ms, is_error, subtype, terminal_reason, permission_denials. Add --json-schema and the validated data lands in structured_output (docs). Parse with jq:

out=$(claude -p "Summarize this diff" --output-format json < diff.patch)
echo "$out" | jq -r '.result'
echo "$out" | jq '{cost: .total_cost_usd, denied: (.permission_denials | length)}'

--output-format stream-json needs --verbose and emits one JSON object per line: hook events, system/init, assistant messages, a rate_limit_event, then a final result. The init event lists mcp_servers, tools and permissionMode, a handy place to assert what the run could touch.

Exit codes, as observed:

Situation Exit Where the error is
Normal run 0 result
Unknown flag 1 stderr: error: unknown option '--bogus-flag'
Unknown model 1 stdout JSON, see below
--max-turns 1 on a task needing a tool 1 stdout JSON, see below
Tool denied by permissions 0 permission_denials array
timeout 3 claude -p ... 124 killed from outside

Raw output for the two JSON rows (--model not-a-model, and --max-turns 1 with a prompt that needs Bash), picked from the --output-format json object:

unknown model  rc=1  {'subtype': 'success', 'is_error': True, 'terminal_reason': 'api_error', 'result': "There's an issue with the selected model (not-a-model). It may not exist or you may not have access to it. ..."}
max-turns 1    rc=1  {'subtype': 'error_max_turns', 'is_error': True, 'terminal_reason': 'max_turns', 'result': None, 'num_turns': 2}

An unknown model reports subtype: "success" next to is_error: true. Gate on is_error and the exit code, not subtype. A run that couldn't do its job because of a denial still exits 0, so in CI also fail on a non-empty permission_denials.

Generating files: print to stdout or let the model Write?

A chore that produces a document has two options: have claude -p print it and redirect stdout, or give it Write and a path. I ran both once on Haiku with the same brief (a 4,000-word PostgreSQL indexing guide), --output-format json, empty strict MCP:

stdout (--tools "") Write (--tools Write, dontAsk)
Exit code 0 0
Wall time 70 s 65 s
Turns 1 2
Output tokens 5,209 5,163
total_cost_usd $0.052 $0.065
Words delivered 3,198 in result 2,776 in guide.md, result was just DONE
stop_reason end_turn end_turn

Nothing was truncated at this size, and both fell short of 4,000 words (single runs, so don't read anything into 3,198 versus 2,776). The difference is in the stdout copy: it opened with I'll write a comprehensive practical guide to PostgreSQL indexing covering all your requested topics. and a --- line before the title. Redirect that to a .md file and the chatter is now your first line. The Write file began at # A Practical Guide.... The price is one extra turn and about 25% more cost here.

Decision rule: if the output is a file that other tools consume, give the model Write and a fixed path (and check the file exists afterwards), because stdout carries preamble you would have to strip. If it's a short answer a script reads, such as a verdict, a label or a commit message, use stdout with --json-schema so the shape is validated. I didn't test outputs near the model's output limit, so I can't tell you where truncation starts.

Claude Code headless timeouts: parallel runs were slower

A trivial say ok on Haiku took 11-18 s solo against about 6 s of API time (duration_api_ms); the rest is startup (settings, hooks, plugins, MCP servers). Then I varied concurrency with conc.sh (default config, N processes launched together):

N=1: 13.7 s
N=2: 14.2 15.1 s
N=4: 15.9 21.5 21.8 23.2 s
N=6: 24.5 26.3 27.3 27.5 28.5 29.7 s

These are single runs per N on one laptop, so treat the trend as indicative. Per-process wall time climbs with N: the median is about 14 s at N=2, 21 s at N=4 and 27 s at N=6. Repeat N=6 runs landed between 20 and 31 s. An earlier N=6 set measured 37.7 to 54.6 s, which contradicts these; I never found out why (I didn't record what else was running), so I'm treating it as an unexplained outlier and not using it. Budget for swings of 2x, not several seconds.

For memory I sampled two ways during N=6 runs (machine total 32,601 MB). System-wide Win32_OperatingSystem.FreePhysicalMemory, once a second, fell from 19,313 MB to a minimum of 17,041 MB: a 2.3 GB drop with other apps open. Then procmem.ps1 polled every 500 ms and summed WorkingSet64 of every process that didn't exist before the run:

22.5 23.4 23.8 24.6 25.1 31.0
peak new-process working set: 3,833 MB total, 1,675 MB in claude.exe

Per run that is about 280 MB in claude.exe and about 640 MB counting every child process (hooks, MCP servers). Working set counts shared pages in each process, so it overstates; the free-memory drop (about 380 MB per run) is a system-wide sample. Budget between the two. I didn't sample CPU, so I can't say whether the slowdown is CPU, disk or API-side. Practical consequences:

  • Set an outer timeout (timeout 600 claude -p ..., exit 124 on expiry) sized to your slow case, not the median.
  • Add --max-turns and --max-budget-usd as inner brakes.
  • Don't launch N runs with & and expect N-times throughput. Cap concurrency at 2-3.
  • Per the docs, a dev server the model starts is killed about 5 s after the result; a background subagent keeps -p open until it finishes or a 10-minute idle ceiling hits.

--strict-mcp-config: skip the cold start

My user config has plugin and claude.ai MCP servers. A default run's init event listed 25 of them (connected, pending, needs-auth, failed) and 171 tools. With --strict-mcp-config --mcp-config '{"mcpServers":{}}' it listed no servers and 32 tools. The CLI reference says the flag means "Only use MCP servers from --mcp-config, ignoring all other MCP configurations."

Same say ok prompt, Haiku, three runs per config, interleaved back to back, same working directory:

Config Wall time (3 runs) Median Cost (3 runs)
Default 12.3, 9.0, 11.9 s 11.9 s $0.042, $0.005, $0.005
--strict-mcp-config (empty) 6.1, 12.7, 6.8 s 6.8 s $0.033, $0.007, $0.004
--strict-mcp-config + --safe-mode 3.5, 3.4, 3.7 s 3.5 s $0.009, $0.002, $0.003

Run 1 of each config paid for cache creation, so compare runs 2 and 3: all three configs are fractions of a cent. The wall-time gap is the more reliable effect, though strict run 2 (12.7 s) shows it's noisy. --safe-mode also disables CLAUDE.md, skills, hooks and plugins, so it's a troubleshooting flag, not a production one.

Two things the docs add that bite in scripts:

  • With --mcp-config and -p, Claude waits for still-pending servers before the first turn, up to MCP_TIMEOUT (30 s by default, v2.1.221+). A slow server adds that wait to every run.
  • Config entries that fail validation are skipped, and the run still exits 0. I passed {"bad":{"url":"http://localhost:9/mcp"}} and system/init reported "mcp_servers":[] and "mcp_server_errors":[{"name":"bad","type":"url_missing_type","message":"Skipped — MCP server \"bad\" has a \"url\" but no \"type\"; add \"type\": \"http\" (or \"sse\" / \"ws\") to this entry"}]. If a job needs an MCP server, fail it when mcp_server_errors is non-empty (v2.1.219+).

A nightly review that has something to review

My first version reviewed git diff --staged. At 02:00 nothing is staged, so a cron job running it reviews an empty diff. A scheduled job needs input that exists without you: here, whatever landed on origin/main in the last 24 hours.

I built a bare origin, a commit dated three days ago, and a second commit from now (get_page made 1-based, plus a new last_page). Then I cloned a "server" copy that had only seen the old commit and ran this against it:

#!/usr/bin/env bash
# usage: nightly.sh /path/to/repo   -- review what landed on origin/main in the last 24 h
set -u
cd "$1" || exit 1
git fetch -q origin || { echo "git fetch failed"; exit 1; }
base=$(git rev-list -1 --before='24 hours ago' origin/main)
[ -n "$base" ] || base=$(git hash-object -t tree /dev/null)   # young repo: review everything
diff=$(git diff "$base" origin/main) || { echo "git diff failed"; exit 1; }
[ -n "$diff" ] || { echo "$(date -Is) nothing landed in 24h"; exit 0; }
out=$(printf '%s\n' "$diff" | timeout 600 claude -p "Review this diff for bugs. Report file:line and one sentence each." \
  --model haiku \
  --tools "Read,Grep,Glob" \
  --permission-mode dontAsk \
  --strict-mcp-config --mcp-config '{"mcpServers":{}}' \
  --max-turns 15 --max-budget-usd 1.00 \
  --no-session-persistence \
  --output-format json)
rc=$?   # status of the pipeline's last command (timeout/claude); printf's is not seen
[ $rc -ne 0 ] && { echo "claude failed rc=$rc"; exit $rc; }
echo "$out" | jq -e '.is_error == false and (.permission_denials | length) == 0' >/dev/null \
  || { echo "$out" | jq -r '.result'; exit 1; }
echo "$(date -Is) review of $base..origin/main"
echo "$out" | jq -r '.result'

rc=$? holds the exit status of the last command in the pipeline (timeout ... claude), so a failure of printf wouldn't show up there; that's why git diff is checked separately, where its failure is reported instead of silently reviewing an empty string.

Why is --tools "Read,Grep,Glob" enough, given the MCP finding above? The empty strict config removes MCP, and --tools removes everything else that isn't listed, Agent included. The init event of a run with exactly these flags showed "tools":["Glob","Grep","Read"] and "mcp_servers":[]: nothing left that can run a command or spawn a subagent.

There is no cron daemon in Git Bash on Windows, so I simulated its environment: env -i with a bare PATH, the cron command line from below (>> review.log 2>&1), from /. An earlier attempt with git missing from that PATH failed the way cron jobs do:

nightly.sh: line 5: git: command not found
git fetch failed

Cron's default PATH is minimal; set it in the crontab. With git, claude and jq on the path, the script exited 0 in about 26 s and review.log ended with:

2026-10-07T03:31:19+05:00 review of 30382599eaa2728a72d1ed8a00d206ba6c0893da..origin/main
**pager.py:8** — `last_page()` uses `page_count()` which only counts complete pages via integer division, so it returns an incomplete page when items aren't evenly divisible by `page_size` (e.g., 10 items with page_size 3 has 4 pages but `page_count()` returns 3).

That's the planted bug: last_page calls the existing floor-division page_count. An earlier run also flagged the pre-existing page_count line, so treat line numbers as pointers, not citations. The log also held two SessionEnd hook ... failed lines (node: command not found under the stripped PATH). They come from my personal user hooks, which -p runs without --bare: a hook in your environment can write into a "clean" log.

The crontab line:

chmod +x nightly.sh
# crontab -e: 02:00 daily, append output to review.log
PATH=/usr/local/bin:/usr/bin:/bin
0 2 * * * /path/to/nightly.sh /path/to/repo >> /path/to/review.log 2>&1

What to do next

  1. Run matrix.sh against your own model and config. It takes about ten minutes and shows which routes to other tools are open in your setup (MCP servers and subagents were the surprise in mine).
  2. Gate CI on the output: fail the job when is_error is true, permission_denials is non-empty, or (if the job needs MCP) mcp_server_errors is non-empty.
  3. Pin a timeout and budget: timeout 600 around the call, timeout-minutes on the CI job, --max-turns plus --max-budget-usd inside, and RAM headroom of roughly 300 to 650 MB per concurrent run (the two measurements above).
  4. Prefer --bare for scripts once you have an ANTHROPIC_API_KEY. The docs call it the recommended mode for scripted and SDK calls; it skips hooks, plugins, MCP discovery and CLAUDE.md.

More on the tooling around chores like these: 5 Essential Tools for Every Full-Stack Developer and Real-World Examples of AI Integration in Web Development. The full list is in the AI Tutorials category.

Not tested: --bare (no API key here), --restricted (v2.1.248+; claude --help says it removes Bash, PowerShell, REPL and other code-running tools, plus WebFetch), SIGTERM exit 143 (docs-only), cron mail behaviour, a real cron daemon, GitHub Actions (see the docs' GitHub Actions page), and generation near the output-token limit.

Bekzod Erkinov

Bekzod Erkinov

Author

Founder of NextGenBeing. Software engineer working with Laravel, Python, and cloud infrastructure. Writes about patterns that actually hold up in production. Based in Tashkent, Uzbekistan.

🎁 Free guide

Get the AI-Assisted Developer's Field Guide

The workflow, prompts, and tools I use to ship faster with AI — free when you subscribe. Plus new deep-dives in your inbox. No spam, unsubscribe anytime.

Comments (0)

Please log in to leave a comment.

Log In

Related Articles

Don't miss the next deep dive

Get one well-researched tutorial in your inbox each week. No spam, unsubscribe anytime.