## Why
Auto Review should remain the effective approval reviewer when settings
cross runtime boundaries. A config or app-server round trip must not
change the reviewer identity, and delegated work must not silently fall
back to user review.
This requires both a stable canonical serialized value and propagation
of the effective setting. `auto_review` is the canonical value across
protocol and app-server output, while `guardian_subagent` remains
accepted as backward-compatible input.
## What changed
- serialize `ApprovalsReviewer::AutoReview` consistently as
`auto_review` across core protocol and app-server v2
- continue accepting `guardian_subagent` when reading existing config or
client requests
- carry the active turn's approval reviewer into spawned agents
- update config/debug expectations and add delegated-task regression
coverage
## Scope
This does not change Guardian policy or remove compatibility with
existing `guardian_subagent` inputs. It preserves the selected reviewer
across serialization, config reloads, app-server settings, and delegated
task setup.
Related Guardian changes are split independently:
- #26231 adds denials and soft denials
- #26334 retries transient reviewer failures
- #26333 reuses narrowly scoped low-risk approvals
- #26232 adds TUI denial recovery
## Validation
- `just test -p codex-app-server-protocol` (224 passed)
- regression coverage for delegated task reviewer propagation
- serialization coverage for canonical `auto_review` output and legacy
`guardian_subagent` input
---------
Co-authored-by: saud-oai <saud@openai.com>
## Why
These workflows currently hard-code the `codex` runner group and custom
runner labels. That makes the same workflow definitions less portable
across repository copies or renamed repos, even though the runner fleet
follows the repository name scheme. Template the runner identities from
the repository name so `openai/codex` still resolves to the existing
`codex-*` runners while other repos can use their own `<repo>-*` runner
names.
## What Changed
- Replaced custom runner `group` values such as `codex-runners` with
`${{ github.event.repository.name }}-runners`.
- Replaced custom runner labels such as `codex-linux-x64` and
`codex-windows-arm64` with `${{ github.event.repository.name }}-...`.
- Covered direct `runs-on` objects, matrix `runs_on` entries, reusable
workflow runner inputs, and release runner labels.
## Verification
- Parsed all `.github/workflows/*.yml` files as YAML with Ruby.
- Searched `.github/workflows` to confirm no hardcoded runner-field
`codex-runners` or `codex-*` labels remain.
## Why
Importing large external-agent session histories currently starts a full
live Codex thread for every imported session. This initializes unrelated
runtime systems and repeats expensive transcript, metadata, hashing, and
ledger work.
On a 50-session, 238 MiB fixture, the existing path took roughly 70
seconds to complete the import and 77 seconds end to end.
## What changed
- Persist imported sessions directly through `ThreadStore` instead of
starting full live threads.
- Process imports through a bounded five-session pipeline.
- Parse, extract, and hash each source file in one pass.
- Move blocking source preparation onto the blocking thread pool.
- Reuse prepared content hashes and update the import ledger once per
batch.
- Avoid metadata readback for newly written rollouts.
- Preserve imported conversation history and visible thread metadata.
- Keep the implementation out of `codex-core` and avoid changes to the
public `ThreadStore` trait.
## Performance
For the same 50-session, 238 MiB fixture:
| Path | Import completion | End to end |
| --- | ---: | ---: |
| Existing import | 69.61s | 76.62s |
| This change | 5.95s | 6.58s |
All 50 sessions imported successfully with no warnings or contention
signals.
## Validation
- `just test -p codex-external-agent-sessions`
- `just test -p codex-app-server external_agent_config_import`
- Verified imports do not initialize unrelated required MCP servers.
- Verified previously imported source versions are skipped and changed
sources can be imported again.
- Verified imported rollouts remain readable through thread listing and
history APIs.
## Why
Compaction analytics adds retained image count and compaction summary
output tokens for v1.5 specifically.
## What changed
- Add nullable `retained_image_count` and `compaction_summary_tokens`
fields to `codex_compaction_event`.
- Populate them only for `responses_compaction_v2`: retained images come
from the retained v2 compacted history, and summary tokens come from
`response.completed.token_usage.output_tokens`.
- Leave local and legacy remote compaction events as `null` for these
detail fields.
## Verification
- `just fmt`
- `just fix -p codex-core`
- `just test -p codex-core
build_v2_compacted_history_counts_retained_input_images`
- `git diff --check`
## Summary
- Keep the existing `x-codex-window-id` HTTP header unchanged.
- Also send the same window ID in Responses `client_metadata`, allowing
supported backend paths to surface it as
`x-client-meta-x-codex-window-id`.
- Cover normal HTTP Responses and remote compaction v2 requests without
changing window generation or compaction behavior.
## Why
In the `2026-06-06T23` production hour, all 28,729 HTTP compaction
requests had `window_id` in `x-codex-turn-metadata`, but only 73
retained the direct `x-codex-window-id` header. The request-body
`client_metadata` path is already used for installation ID and is
preserved through supported Responses API paths.
This is additive metadata only. It does not change the direct header,
request count, model input, compaction routing, window generation, or
user response behavior.
Legacy `/v1/responses/compact` is intentionally unchanged. Its current
server-side `CompressBody` schema does not accept `client_metadata` and
rejects unknown fields, so supporting that path requires a backend
schema change before the Codex client can safely send this field.
## Validation
- Current head: `219baef3c`, rebased onto `origin/main` at `26d932983`.
- The post-rebase diff remains limited to the original five files (`22`
insertions, `6` deletions); the legacy experiment remains fully
reverted.
- `just test -p codex-core
responses_stream_includes_subagent_header_on_review`: passed; validates
normal HTTP Responses metadata.
- `just test -p codex-core
remote_compact_v2_reuses_compaction_trigger_for_followups`: passed;
validates remote compaction v2.
- `just test -p codex-core
remote_manual_compact_chatgpt_auth_reuses_service_tier_and_prompt_cache_key`:
passed; validates that legacy compact keeps its accepted payload shape.
- `just test -p codex-core
remote_manual_compact_api_auth_omits_service_tier_and_reuses_prompt_cache_key`:
passed; validates the legacy API-key payload as well.
- `just fmt`: passed; an unrelated root `justfile` rewrite produced by
the formatter was discarded.
- `git diff --check origin/main...HEAD`: passed.
The focused server pytest could not start in the local monorepo
environment because test setup is missing the `dotenv` module. Server
source and tests explicitly show that `CompressBody` omits
`client_metadata` and `/v1/responses/compact` returns HTTP 400 for
unknown body fields.
## Why
Remote-control app-server sessions can reconnect every 5-7 seconds when
the shared transport-event queue fills. The queue's consumer handled
`ConnectionClosed` by awaiting all in-flight RPCs for the disconnected
connection. A stuck RPC therefore blocked processing of replacement
connection and initialize events until remote-control forwarding hit its
five-second timeout and reconnected again.
Related issue: N/A (internal remote-control incident investigation).
## What Changed
- Split fast RPC admission closure from draining:
`ConnectionRpcGate::close()` rejects queued and future RPCs, while
`shutdown()` continues waiting for RPCs that already started.
- Close a disconnected connection's RPC gate before spawning the
existing RPC drain and resource cleanup in a tracked background task, so
the transport-event consumer remains available without waiting for
active RPCs.
- Reap completed cleanup tasks during normal operation, drain them
during graceful shutdown, and abort them during forced shutdown.
- Add regression coverage for closing with an active RPC, rejecting
post-close requests without polling them, and preserving the existing
shutdown wait behavior.
## Verification
`just test -p codex-app-server --lib connection_rpc_gate` passes all 6
tests, including the new close-versus-drain regression coverage.
## Summary
- Restore separate release symbol archives for macOS, Linux, and Windows
binaries.
- Build release binaries with `line-tables-only` debuginfo instead of
full debuginfo.
- Strip Unix distribution binaries after extracting symbols, preserve
Windows PDBs, and keep symbol archives available to the release job.
- Strip the packaged Linux `bwrap` binary before hashing it so the
embedded digest matches the distributed bytes.
## Root cause
The first symbol-artifact implementation enabled
`CARGO_PROFILE_RELEASE_DEBUG=full`. In the June 2 release runs, macOS
ARM primary builds reached the 90-minute timeout while still inside
`Cargo build`. After the symbol changes were reverted, the same primary
build completed in about 22 minutes. The archive step itself completed
in tens of seconds when reached.
Rust's `line-tables-only` debuginfo level preserves function names and
source locations for symbolication without emitting the heavier variable
and type information from full debuginfo.
## Validation
- Ran `just fmt` from `codex-rs`.
- Ran `just test-github-scripts` from the repository root: 23 tests
passed.
- Ran `bash -n` and `shellcheck` on
`.github/scripts/archive-release-symbols-and-strip-binaries.sh`.
- Parsed both modified workflows as YAML and ran `git diff --check`.
- Built a macOS release smoke binary with `line-tables-only`, archived
its dSYM through the restored script, stripped the production binary,
and verified that `atos` resolves `symbol_smoke_function` to
`main.rs:2`.
- Ran Linux archive-script control-flow coverage with stubbed `objcopy`
and `strip` commands.
- Ran Windows PDB archive staging coverage and verified
underscore-emitted Rust PDB names are staged under shipped hyphenated
binary names.
## Follow-up
The release workflow only runs for tags or manual dispatches, so CI
cannot dry-run the full release matrix on this PR. The next release run
will verify runner time and memory behavior under `line-tables-only`.
## Summary
- add contains_external_context() to tool output so other tools can be
opted out of influencing memory when disable_on_external_context=true
- Classify standalone web-search output as external context (to match
behavior as hosted web search)
- Verify with integration test
## Why
Multi-agent v2 residency is intended to keep only the threads that need
to be live. The existing rollout resume path still walked persisted open
descendants and reopened the entire descendant tree when resuming a v2
root, which turns resume into an eager reload of work that should stay
unloaded until it is explicitly needed.
The interrupted-agent path has a related residency issue. Interrupted
agents remain open by design, so an idle interrupted resident should be
eligible for eviction just like an idle completed or errored resident.
Otherwise a resident set full of interrupted agents can consume every v2
slot and block later spawns or reloads with `AgentLimitReached`.
## What Changed
- Return early from `resume_agent_from_rollout` after resuming a v2
thread so persisted v2 descendants are not reopened eagerly.
- Treat idle `Interrupted` v2 residents as unloadable in the LRU
residency path.
- Add focused coverage for v2 root resume leaving descendants unloaded
and for eviction of an idle interrupted v2 resident when a new slot is
needed.
## Verification
Added targeted `codex-core` tests covering:
- v2 root resume with persisted descendants, verifying only the root is
loaded after resume.
- residency eviction of an idle interrupted v2 agent when the resident
set is full.
## Why
`close_agent` is the wrong model-facing name for the v2 operation after
the residency changes. V2 agents remain reusable by task name, and
residency/unloading owns capacity management; the exposed tool should
describe the action it actually performs: interrupt the target agent's
current turn without making the agent unavailable for future messages or
follow-up tasks.
## What changed
- Rename the multi-agent v2 tool from `close_agent` to
`interrupt_agent`.
- Keep the v1 `close_agent` surface unchanged.
- Update the v2 handler to send `Op::Interrupt`, keep interrupted agents
registered, and reject root/self targets with interrupt-specific errors.
- Route interrupt delivery through the existing dead-thread cleanup path
so stale resident entries do not keep consuming capacity.
- Update tool planning and handler tests for the new v2 surface and
semantics.
## Verification
Added focused coverage in:
- `core/src/tools/spec_plan_tests.rs`
- `core/src/tools/handlers/multi_agents_tests.rs`
## Why
Multi-Agent V2 concurrency should count active non-root turns, not
resident or durable agent threads. The limit is intentionally best
effort: admission checks are synchronous, but concurrent successful
checks may overshoot slightly.
## What changed
- Keep one root-derived execution limit on the shared `AgentControl`.
- Count active V2 subagent turns with an RAII guard owned by
`RunningTask`.
- Check capacity before spawning or starting an idle agent, including
direct app-server `turn/start` submissions.
- Preserve queued delivery for agents that are already running.
- Exempt automatic idle continuations so `/goal` work is not dropped
when capacity is temporarily full.
- Keep root and V1 turns outside this limiter.
## Test coverage
- `execution_guards_count_active_v2_subagent_turns`
- `execution_guards_ignore_root_and_v1_turns`
- `v2_nested_spawn_checks_shared_active_execution_capacity`
## Summary
- ignore RUSTSEC-2026-0173 in cargo-deny and cargo-audit config
- document that proc-macro-error2 is pulled in transitively via
i18n-embed-fl/age/codex-secrets
- leave the ignore temporary until codex-secrets moves off age or age
drops i18n-embed-fl
## Validation
- just fmt
- cargo deny check --hide-inclusion-graph
## Why
Multi-agent v2 treats agents as durable logical agents, not just live
entries in `ThreadManager`. After the reload-on-delivery change, a v2
agent can be addressed even if its thread is not currently loaded.
This PR adds the next layer: loaded v2 subagents can be paged out of
`ThreadManager` when the session has too many resident agents. That
keeps residency separate from logical identity and prepares the stack
for making v2 concurrency count active execution instead of existing
agents.
## What Changed
- Add an `AgentControl`-scoped LRU for resident v2 subagents.
- Reserve residency before spawning or reloading a v2 subagent.
- If resident capacity is full, unload the least-recently-used idle v2
subagent from `ThreadManager`.
- Keep `ThreadManager` as a primitive loaded-thread store; it does not
own the LRU policy.
- Keep unloaded agents registered and durable so they can be reloaded by
the delivery path.
- Preserve the existing v2 cap semantics by using the derived non-root
v2 cap for residency.
Eviction is intentionally conservative. A thread is unloadable only when
it is a v2 subagent, has completed or errored, has no active turn, and
has no pending mailbox work. Before removal, the rollout is materialized
and flushed.
## Assumptions And Non-Goals
- PR #26623 provides the reload-on-delivery path for unloaded v2 agents.
- `ThreadManager` membership means loaded/resident, not logical agent
existence.
- `AgentRegistry` remains the logical identity/metadata source for v2
agents that may be unloaded.
- `list_agents` remains a recent/resident view for now.
- This does not change active execution concurrency; that is the next
PR.
- This does not change `close_agent` semantics.
- This does not change or remove `resume_agent`.
- This does not add a new residency config knob.
## Stack
1. V2 durable lookup and reload on delivery (#26623) - reload unloaded
v2 agents before delivering follow-up/input.
2. V2 residency LRU (this PR) - unload idle resident v2 agents from
`ThreadManager` when resident capacity is full.
3. V2 active-execution concurrency - count running non-root v2 turns
instead of logical agents.
4. V2 close/interrupt semantics - make v2 close interrupt the current
turn without deleting durable identity.
5. V2 resume cleanup - remove the manual resume surface for v2 while
keeping internal reload support.
## Validation
- Added focused coverage for the residency LRU eviction path.
- Local clippy/check/tests were not run; CI will cover them.
## What
- Consume plaintext `output` from standalone search while retaining
optional `encrypted_output` parsing.
- Expose `web.run` to code mode and return search output to nested
JavaScript calls.
- Cover direct and code-mode standalone search paths with integration
tests.
## Why
`/v1/alpha/search` now returns plaintext output, which code mode needs
to consume standalone search results.
## Test plan
- `just test -p codex-api`
- `just test -p codex-web-search-extension`
- `just test -p codex-core code_mode_can_call_standalone_web_search`
- `just test -p codex-app-server
standalone_web_search_round_trips_output`
## Why
MCP startup failures from spawned subagents were rendered as global
notifications, so a child thread's failure could pollute the visible
parent transcript. Routing the notification to the child exposed two
related replay problems: session refresh could discard the buffered
event, and a newly created child `ChatWidget` did not know the expected
MCP server set, which could leave its startup spinner running after
every server had settled.
MCP startup diagnostics should remain visible in the thread that owns
the startup without affecting other transcripts. The protocol also needs
to support a future app-scoped MCP lifecycle where startup is not owned
by any thread.
## Reported Behavior
The [originating Slack
report](https://openai.slack.com/archives/C08JZTV654K/p1780604538859939)
called out that using subagents could turn MCP startup failures into a
wall of yellow CLI warnings because repeated failures were not
deduplicated. The intended behavior is for those diagnostics to remain
visible once in the thread that owns the startup, without polluting the
parent transcript.
## What Changed
- add nullable `threadId` ownership to `mcpServer/startupStatus/updated`
- populate it from the app-server conversation ID for the current
thread-scoped lifecycle and regenerate the protocol schema and
TypeScript artifacts
- treat a missing or null `threadId` as app-scoped without injecting it
into the active chat transcript
- route and buffer thread-owned MCP startup notifications by thread in
the TUI
- preserve buffered MCP startup events across child session refresh
- seed expected MCP servers before replaying a thread snapshot so
startup reaches its terminal state
- suppress an identical repeated failure warning for the same server
within one startup round
The owning thread still renders the detailed failure and final `MCP
startup incomplete (...)` summary.
## How to Test
1. Configure an optional MCP server named `smoke` that exits during
initialization.
2. Launch the TUI with multi-agent support enabled.
3. Confirm the main thread's own startup failure renders one detailed
`smoke` warning and one incomplete-startup summary.
4. Spawn exactly one subagent.
5. Confirm the parent transcript does not receive the subagent's MCP
startup failure.
6. Switch to the subagent thread and confirm it contains exactly one
detailed `smoke` failure and one incomplete-startup summary.
7. Confirm the subagent's MCP startup spinner disappears and the thread
remains usable.
8. Switch between the parent and subagent and confirm the warnings
neither move nor duplicate.
Targeted tests:
- `just test -p codex-app-server-protocol`
- `just test -p codex-app-server
thread_start_emits_mcp_server_status_updated_notifications`
- `just test -p codex-tui mcp_startup`
The parent/child behavior and spinner completion were also exercised
manually in tmux. `just argument-comment-lint` was attempted but blocked
by an unrelated local Bazel LLVM empty-glob failure; touched Rust
callsites were inspected manually.