Codex Hooks Make the Harness Real

From the guide: Codex CLI Comprehensive Guide

As of Codex 0.150.1 (August 27, 2026), Codex hooks register twelve lifecycle events, refuse to run any non-managed hook until you review and trust its exact definition, and ship enabled by default. The feature that shipped to general availability in the May 14 launch bundle alongside the ChatGPT mobile app has matured into a governance surface: PreToolUse can block or rewrite a tool call before it runs, PermissionRequest can decide an approval, PostToolUse can replace the result the model sees, and Stop can refuse to let a turn finish.23

Codex no longer looks like a coding assistant that waits inside one terminal. It looks like an operating layer that follows work across machines, approvals, projects, chats, diffs, tests, screenshots, plugins, credentials, and local tools.4

Codex hooks make the harness real. Once the agent can work from a phone, reach remote development environments, and run lifecycle hooks, teams need a control system around the model: evidence, approvals, git custody, source discipline, and taste.

TL;DR

Codex supports the workflow shape that agent teams have been building privately: long-running work, remote execution, mobile steering, approvals, hooks, scoped credentials, and audit signals.245 The hooks engine at 0.150.0 registers twelve events, every non-managed hook stays skipped until you trust its current hash, and hooks load from hooks.json files or inline [hooks] tables in config.toml.37 The practical question is not “how do we prompt Codex?” The practical question is “what must Codex prove before we trust the result?” Teams should use hooks and configuration to encode review gates, security boundaries, public-writing standards, and release discipline. They should keep private machinery private and publish only the pattern, the acceptance criteria, and the verified outcome.

Key Takeaways

For engineering teams:

  • Treat Codex hooks as process infrastructure, not decoration. The trust-review flow is part of that infrastructure, not friction to bypass.
  • Start with evidence, approvals, git custody, and release checks before adding clever automation.

For agent-tool builders:

  • Build around Codex’s real surfaces: mobile control, Remote SSH hosts, sandbox modes, approval policies, project instructions, hooks, telemetry, and version control.
  • Port jobs-to-be-done, not old slash-command shapes.

For public writers:

  • Use the official docs at learn.chatgpt.com for current Codex behavior, and check the engine source when the docs lag a release.
  • Describe private practice as author analysis, and leave private prompts, hook bodies, file paths, source lists, credentials, and scoring internals out of public copy.

Where Did Codex Hooks Come From?

OpenAI published “Work with Codex from anywhere” on May 14, 2026.1 The docs changelog entry for that date records the launch bundle: Codex became usable from the ChatGPT mobile app by connecting it to a Mac running the Codex app, hooks reached general availability, and Codex access tokens arrived for trusted automation.2 Codex runs from the connected host, so the same projects, files, credentials, plugins, skills, and configuration are available from a phone.2

Remote connections extend the reach past one desk. The announcement’s Remote SSH capability lands in the docs as SSH hosts: the ChatGPT desktop app can add remote projects from an SSH host and run chats against the remote filesystem and shell. The framing is concrete: “Remote access uses the connected host’s projects, chats, files, credentials, permissions, plugins, Computer Use, browser setup, and local tools.”4

Hooks themselves predate the launch as an experiment and outgrew it afterward. The docs define them as an extensibility framework that runs scripts or MCP tools during the agentic loop, and name plain jobs: send chats to a logging engine, block accidentally pasted API keys, summarize chats into persistent memories, run a validation check when a turn stops, and customize prompting per directory.3 Hooks are enabled by default now; features.hooks in config.toml works as a kill switch, and features.codex_hooks survives only as a deprecated alias.36

Those details matter because they turn agent work from a chat exchange into governed operations.

What Do Codex Hooks Look Like In Configuration?

A hooks post should show a hook. Codex discovers hooks next to active config layers, most usefully ~/.codex/hooks.json, <repo>/.codex/hooks.json, or inline tables in either layer’s config.toml; when several sources exist, all matching hooks load and run.3 A minimal hooks.json with a tool gate and a completion gate:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "^Bash$",
        "hooks": [
          {
            "type": "command",
            "command": "python3 ~/.codex/hooks/pre_tool_use_policy.py",
            "timeout": 30,
            "statusMessage": "Checking Bash command"
          }
        ]
      }
    ],
    "Stop": [
      {
        "hooks": [
          {
            "type": "command",
            "command": "python3 ~/.codex/hooks/evidence_gate.py"
          }
        ]
      }
    ]
  }
}

The same shape written inline in config.toml:

[[hooks.PreToolUse]]
matcher = "^Bash$"

[[hooks.PreToolUse.hooks]]
type = "command"
command = "python3 ~/.codex/hooks/pre_tool_use_policy.py"
timeout = 30
statusMessage = "Checking Bash command"

Three levels organize every hook: an event, a matcher group that decides when the event applies, and one or more handlers (command or mcp_tool).3 The current events, with the engine’s twelve-entry list as the authority:37

Event Fires when Can block?
SessionStart A session starts: startup, resume, clear, or compact; stdout becomes developer context Yes: continue: false stops the hook run, and after compaction it ends the turn
UserPromptSubmit Before a user prompt reaches the model Yes: decision: "block" rejects the prompt
PreToolUse Before a supported tool call runs Yes: deny the call, or rewrite it with updatedInput
PermissionRequest Codex is about to ask for approval Yes: allow or deny; silence falls through to the normal prompt
PostToolUse After a supported tool produces output, including failed commands Partly: replaces the result, cannot undo side effects
PreCompact Before Codex compacts the chat, manual or auto Yes: continue: false stops compaction
PostCompact After Codex compacts the chat Yes: continue: false stops after compacting
SubagentStart A subagent starts, matched on agent_type No: continue: false is parsed but does not stop the subagent
SubagentStop A subagent stops Yes: decision: "block" sends the subagent back for another pass
Stop A turn tries to finish Yes: decision: "block" keeps Codex going, with your reason as the continuation prompt
SessionEnd The main thread ends; never for subagents No: advisory only, 1-second default timeout with a 3-second cap
Interrupt An active top-level turn is interrupted (0.150.0+); never for subagents No: informational, same 1-second default and 3-second cap as SessionEnd

The live hooks doc still documents eleven of those events; it has no Interrupt section yet. The 0.150.0 changelog entry (#40511) and HOOK_EVENT_NAMES: [&str; 12] in the engine source at rust-v0.150.0 carry the twelfth.27 When a page and the engine disagree, trust the engine.

Which Tool Calls Can Hooks Actually See?

Earlier versions of the docs warned that PreToolUse covered little beyond shell and MCP calls. The current doc widens the surface: “PreToolUse and PostToolUse can observe more than shell and MCP calls. Most local function tools use the same hook path,” so a matcher can name tools like update_plan directly, and spawn_agent also matches as Agent.3 Shell commands match as Bash, apply_patch file edits match as apply_patch, Edit, or Write, and MCP tools match names like mcp__filesystem__read_file.3

Hosted tools remain outside: WebSearch and its peers never pass through the local function-tool hook path.3 The doc keeps a softened caveat worth quoting whole: “Some specialized tool paths can opt out of the default hook path. Treat tool hooks as a useful guardrail, not a complete enforcement boundary.”3 Sandboxing still owns the hard boundary; hooks own review and steering inside it.

How Does Codex Decide Which Hooks May Run?

Hooks are code that steers the agent, so Codex governs the hooks themselves. The control system around the model starts here.

Before any non-managed hook runs, Codex requires you to review and trust its exact definition. Trust is recorded against the hook’s current hash, so a new or edited hook is marked for review and skipped until trusted again.3 The /hooks command in the CLI opens the review surface: inspect hook sources, review new or changed hooks, trust them, or disable individual ones. When hooks need review at startup, Codex prints a warning pointing at /hooks.3

Managed hooks from system, MDM, cloud, or requirements.toml sources sit above that flow: they are trusted by policy and cannot be disabled from the user hook browser.3 Plugins sit inside it: installing or enabling a plugin does not trust its bundled hooks, which stay skipped until reviewed like any other.3 Project-local hooks load only when the project’s .codex/ layer is trusted; an untrusted project still loads your user and system hooks.3

Automation feels the same rule. A codex exec run has no review UI, so an untrusted hook is skipped silently until you trust it in an interactive session first; the doc does not say so yet, but the engine’s startup review surface exists only in the interactive TUI.7 For pipelines that vet hook sources somewhere else, --dangerously-bypass-hook-trust runs enabled hooks without persisted trust for that one invocation.3 Trust it deliberately or watch your gate not fire: the harness decides which code may steer the agent before any of it runs.

What Happens When A Hook Blocks, And When Is It Too Late?

Timing decides what a hook can still change. Most hooks run synchronously with a 600-second default timeout; SessionEnd and Interrupt default to one second and are capped at three: SessionEnd fires while the session is tearing down, and Interrupt fires while the user is waiting.37

PreToolUse acts before anything happens, so it holds the strongest cards: deny the call, or rewrite it by returning permissionDecision: "allow" with updatedInput. The deny shape:

{
  "hookSpecificOutput": {
    "hookEventName": "PreToolUse",
    "permissionDecision": "deny",
    "permissionDecisionReason": "Destructive command blocked by hook."
  }
}

Exit code 2 with the reason on stderr blocks too.3

PostToolUse acts after the tool ran, so it cannot undo side effects. A decision: "block" replaces the tool result with your feedback and continues the model from that message, which corrects course without pretending the command never happened.3 Stop turns a refusal into a continuation: block a completion and Codex keeps working, with your reason as the new prompt.3

For checks that should never sit on the critical path, set async = true on a command handler. Background hooks run while Codex continues, deliver output at the next safe point, and explicitly cannot block, approve, or rewrite anything; keep tool policies, permission decisions, prompt rejection, and turn continuation synchronous.3

What Upgraded In 0.149 And 0.150?

Two stable releases in late August tightened the governance story.

Codex CLI 0.149.0 (August 20, 2026) retired the untrusted approval policy (#39630); a config that still names it now fails with an actionable error telling you to remove the setting.27 Codex CLI 0.150.0 (August 26, 2026) added the Interrupt hook event (#40511): “New Interrupt hooks can run commands or MCP handlers when an active top-level turn is interrupted.” Interrupt hooks never run for subagents.27 The same release stopped untrusted projects from supplying project-level AGENTS.md instructions (#39837), which pairs with the existing rule that project-local hooks load only from a trusted .codex/ layer.23

The direction is consistent: an untrusted directory gets less and less authority over the agent that walks into it. Instructions, hooks, and approval shortcuts all now flow through explicit trust decisions.

Why Do Hooks Matter More Than Mobile?

Mobile access changes where the human can intervene. Hooks change what the system can enforce.

A phone lets an operator answer a question while away from the desk. A hook can catch the agent before a risky action, after a file edit, before completion, or during a release check. The phone solves latency. The hook solves standards.

Codex already has first-party control surfaces around sandboxing and approvals. The safety docs pair sandbox mode, which defines what the agent can technically do, with approval policy, which defines when Codex must stop and ask before acting.5 The agent runs with network access turned off by default, and the default local workspace-write mode keeps network access off unless the user enables it.5 Hooks sit next to those controls as the review and steering layer, not a replacement for sandboxing.

Hooks can make local standards executable:

Standard Hook-shaped enforcement
Do not leak secrets Scan prompts and tool inputs (UserPromptSubmit, PreToolUse) before risky actions
Do not fake completion Stop completion (Stop) when evidence is missing
Do not publish stale writing Require source checks and rendered-route checks before release
Do not leave dirty state Require exact-path git status and commit intent (PostToolUse, Stop)
Do not weaken quality Run focused review gates (PermissionRequest, Stop) before release

The model can forget a rule. A hook can re-run the rule at the moment the rule matters.

What Does The Harness Own That The Provider Does Not?

An agent harness is the operating layer around a model: permissions, memory, tools, hooks, source checks, release gates, review packets, and rollback discipline. The term can sound private or ornate, but the job is plain. The layer turns intent into accountable work.

Codex now exposes enough official surface to make that layer explicit. Remote connections carry the host environment. Sandbox modes and approval policies define action boundaries. Configuration files define models, projects, permissions, MCP servers, skills, hooks, telemetry, and features.6 OpenTelemetry export stays opt-in and off by default; when enabled, Codex emits structured events covering chats, API requests, stream activity, user prompts (redacted by default), tool approval decisions, and tool results.58

That set of surfaces creates a useful split:

Provider surface Team-owned standard
Remote connection Which hosts and accounts can carry work
Sandbox and approvals Which actions deserve friction
Hooks Which standards run at decision points
Hook trust Which code may steer the agent at all
Telemetry Which events become audit evidence
Git workflow Which changes become save points
Project instructions Which durable norms guide the agent

The provider should keep improving the runtime. The team still owns judgment.

What Should Teams Encode First?

Start with four gates. They pay rent immediately.

Evidence Gate

Codex’s original launch post emphasized verifiable evidence: terminal logs, test outputs, and traceable steps during task completion.9 Make that expectation non-negotiable. A meaningful completion should name the files changed, commands run, observed behavior, failed checks, and remaining gaps.

For public work, evidence includes source links and claim-source alignment. For web releases, evidence includes rendered routes, metadata, schema, discovery files, deployment state, cache freshness, and live changed markers. For translations, evidence includes locale coverage, quality gates, storage rows or cache files, and native-review status when required.

Approval Gate

Do not use one approval posture for every action. The approvals doc’s current combination table runs from the Auto preset (workspace-write sandbox with on-request approvals) through safe read-only browsing, read-only non-interactive CI, auto-review mode, and dangerous full access.5 One row lags reality: the untrusted policy still appears on the page, but 0.149.0 retired it and explicit configs now error.25 For an always-ask posture today, combine the read-only sandbox with on-request approvals. A strong local policy keeps the same shape: low-risk reads pass quietly, side-effecting work gets review, and destructive or externally visible work gets explicit evidence.

Git Custody Gate

Agent work needs rollback handles. Codex’s own security docs say Codex works best with version control: keep status clean before delegating, commit frequently, run targeted verification, review diffs, and document decisions in commit messages.5

That advice should become process. Commit after coherent, verified save points. Stage exact paths. Split commits by independently revertible concern. Ask before push unless the release flow already grants publishing authority. Do not sweep unrelated dirty files into a commit because the agent happened to see them.

Taste Gate

AI coding makes implementation cheaper. Cheaper implementation raises the value of taste.

Taste does not mean decorative preference. It means the work improves the whole product. It means the agent can refuse a technically possible path that weakens the result. It means public writing avoids private machinery, unsupported claims, and filler. It means a correct local patch can still fail if the user-visible path remains broken.

A taste gate should ask:

Question Purpose
Who is the real user? Prevent local artifact worship
What proves the outcome? Separate evidence from confidence
What did we remove or refuse? Preserve coherence
What remains unverified? Avoid false completion
Why does the work deserve to exist? Keep volume from replacing judgment

What Does Mozilla’s Firefox Work Prove?

Mozilla’s May 7 post about hardening Firefox with Claude Mythos Preview makes the same point from a different stack. The team says early LLM code-audit attempts showed promise but had too many false positives to scale. Agentic harnesses changed the economics because they could create and run reproducible test cases to dynamically test bug hypotheses.10

Mozilla’s important sentence is not about the model alone. The team says discovery was necessary but not sufficient. The useful system had to integrate with the full security bug lifecycle: targets, deduplication, bug tracking, triage, fixes, and release.10 The authors also say the pipeline reflected Firefox’s codebase semantics, tooling, and processes.10

That is the lesson for Codex. Better models matter. The operational system around the model decides whether the work becomes trusted output.

What Should Stay Out Of Public Copy?

A public Codex article should not dump the private working system.

Keep these out of public copy:

  • private prompts and hook bodies;
  • sensitive local paths;
  • exact source maps and scoring internals;
  • account identifiers and credential handling;
  • private workflow shortcuts;
  • unreleased plugin behavior;
  • anything that helps a stranger reconstruct internal operations.

Publish the pattern instead: what the gate protects, what evidence it requires, what failure it catches, and how a team can implement the idea using official Codex surfaces.

That line protects trust. It also improves the writing. Private machinery usually reads like folklore. Public acceptance criteria help other teams reason about their own systems.

What Does A Minimal Codex Harness Map Look Like?

Build the smallest control map that proves useful work.

Layer First useful version
Project policy AGENTS.md with durable norms and verification commands
Permissions Workspace-write by default, explicit network and external writes
Hooks Secret scan, evidence stop gate, git custody, public-writing checks
Hook trust Reviewed hashes; bypass flag only in pipelines that vet sources elsewhere
Source discipline Primary-source verification for current tool behavior
Review packet Goal, changed files, commands, results, sources, gaps
Git custody Exact-path commits after verified save points
Release gate Rendered route, metadata, schema, translations, live markers
Telemetry Approval, tool, and network events routed to trusted collectors

Start explicit. Run one real task. Record where the gate helped and where it got in the way. Promote only the parts that improve the user-visible outcome.

Quick Summary

Codex hooks, Remote SSH, mobile control, sandboxing, approvals, configuration, telemetry, and version control point in the same direction: coding agents need operating systems around them.2456 The agent can write code. The harness decides what counts as work.

The best teams will not win by producing the most agent output. They will win by making agent work inspectable, reversible, sourced, tasteful, and worthy of release.

FAQ

What are Codex hooks?

Codex hooks run scripts or MCP tools during the agentic loop, loaded from hooks.json files or inline [hooks] tables in config.toml. The docs name the jobs plainly: send chats to a logging engine, block accidentally pasted API keys, summarize chats into persistent memories, run validation when a turn stops, and customize prompting per directory.3 The engine at 0.150.0 registers twelve events, from PreToolUse, PermissionRequest, and PostToolUse through Stop and the new Interrupt; the doc page still lists eleven while it catches up.37

Why do Codex hooks matter?

Hooks let teams put standards at decision points instead of relying only on prompts. A hook can check evidence, source quality, git state, or release readiness when the agent acts or tries to finish.

Why did my hook not run?

The usual answer is trust. Codex skips any non-managed hook whose current hash you have not reviewed, prints a startup warning in interactive sessions, and skips silently in codex exec automation.7 Open /hooks to review and trust it, or pass --dangerously-bypass-hook-trust only in pipelines that vet hook sources elsewhere.3

Does Codex mobile replace local agent workflow?

No. Mobile control lets users steer work away from the desk, but the connected host still supplies the projects, chats, files, credentials, permissions, plugins, and local tools.4 Teams still need local policy, safe credentials, version control, and verification.

What should a Codex harness include first?

Start with project instructions, sandbox and approval posture, a secret boundary, an evidence stop gate, exact-path git custody, source verification for public claims, and a release gate for user-visible work.

Should teams publish their Codex hooks?

Publish patterns and acceptance criteria, not private hook bodies or sensitive workflow details. A useful public post can explain the job of a hook without exposing private paths, source maps, prompts, credentials, or scoring rules.

References


  1. OpenAI, “Work with Codex from anywhere,” OpenAI, May 14, 2026. ↩

  2. OpenAI, “ChatGPT & Codex changelog,” ChatGPT Learn, accessed August 28, 2026. May 14, 2026 entry (mobile launch, hooks general availability, Codex access tokens for trusted automation) and Codex CLI 0.149.0, 0.150.0, and 0.150.1 release entries. ↩↩↩↩↩↩↩↩↩↩

  3. OpenAI, “Hooks,” ChatGPT Learn, accessed August 28, 2026. ↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩↩

  4. OpenAI, “Remote connections,” ChatGPT Learn, accessed August 28, 2026. ↩↩↩↩↩

  5. OpenAI, “Agent approvals & security,” ChatGPT Learn, accessed August 28, 2026. ↩↩↩↩↩↩↩↩

  6. OpenAI, “Configuration Reference,” ChatGPT Learn, accessed August 28, 2026. ↩↩↩

  7. openai/codex at rust-v0.150.0, GitHub, accessed August 28, 2026: HOOK_EVENT_NAMES in codex-rs/hooks/src/lib.rs; the timeout normalization in codex-rs/hooks/src/engine/discovery.rs; the Interrupt subagent early-return in codex-rs/core/src/hook_runtime.rs; the startup hook-review surface living only in the tui crate, with untrusted handlers excluded without warning in discovery.rs and --dangerously-bypass-hook-trust global on exec in exec/src/cli.rs. ↩↩↩↩↩↩↩↩↩

  8. OpenAI, “Running Codex safely at OpenAI,” OpenAI, May 8, 2026. ↩

  9. OpenAI, “Introducing Codex,” OpenAI, May 16, 2025. ↩

  10. Brian Grinstead, Christian Holler, and Frederik Braun, “Behind the Scenes Hardening Firefox with Claude Mythos Preview,” Mozilla Hacks, May 7, 2026. ↩↩↩

Related Posts

Agent Skills Need Package Managers

Agent skills, MCP servers, prompts, hooks, and commands now behave like dependencies. Teams need manifests, lockfiles, p…

13 min read

Install and Update Codex CLI: Mac, Linux, Windows

codex update upgrades script, npm, and Homebrew installs; winget upgrade OpenAI.Codex covers winget. Install, update, pi…

19 min read

Two MCP Servers Made Claude Code an iOS Build System

XcodeBuildMCP and Apple's Xcode MCP give Claude Code structured access to iOS builds, tests, and debugging. Setup, real-…

19 min read