Commit Graph

8059 Commits

Author SHA1 Message Date
sayan-oai
b07e13ea61 Narrow 0.144 hotfix to GPT-5.6 prompts and context (#34009)
## Summary

- Keep the refreshed bundled prompts for GPT-5.6 Sol, Terra, and Luna.
- Keep the 272,000-token context and maximum context windows for those
three models.
- Revert the unrelated model-catalog metadata changes introduced by
#33972.

## Why

#33972 backported a broader generated catalog refresh than this stable
hotfix needs. The intended `0.144.6` change is limited to the GPT-5.6
Sol/Terra/Luna prompt refresh and their corrected context-window
metadata.

This follow-up restores all other catalog values to the last-green
`release/0.144` baseline. In particular, it removes fields that `0.144`
does not understand, restores the existing skills-guidance settings and
GPT-5.5 availability NUX, and leaves model visibility and upgrade
behavior unchanged.

Relative to `215fb464f1932b5fd7beb5d4615bc1927d80fbf7`, the resulting
catalog differs in exactly 12 values: `base_instructions`,
`model_messages.instructions_template`, `context_window`, and
`max_context_window` for each of GPT-5.6 Sol, Terra, and Luna.

## Validation

- `just test -p codex-models-manager bundled_models_json_roundtrips`
- `just fmt`
- `jq empty codex-rs/models-manager/models.json`
- `git diff --check`
- Verified the semantic diff against
`215fb464f1932b5fd7beb5d4615bc1927d80fbf7` contains only the intended 12
values.
- Verified all retained values exactly match
`d26a9bf671b1c03aabfc32e1092d137c1feb3962`.
2026-07-18 05:49:28 -07:00
sayan-oai
9cc6a1ee4c Backport refreshed bundled model metadata to 0.144 (#33972)
## Summary

- Backport `d26a9bf671b1c03aabfc32e1092d137c1feb3962` to the `0.144`
release line.
- Refresh bundled GPT-5.6 model instructions and context-window
metadata.
- Refresh reasoning-summary, skills, permissions, and auto-review
catalog metadata.

## Why

The catalog refresh has propagated to public `main`, but stable `0.144`
clients still bundle the older model metadata. This backport makes the
refresh available in the next non-alpha `0.144.x` hotfix.

The release branch's existing `supports_reasoning_summaries` values were
preserved while adding `supports_reasoning_summary_parameter` from the
source commit.

## Validation

- `just test -p codex-models-manager` (40 passed)
- `jq empty codex-rs/models-manager/models.json`
- Verified GPT-5.6 Sol does not include the `ultrafast` service tier
2026-07-18 04:45:29 -07:00
Dylan Hurd
215fb464f1 [release/0.144] fix(core) expand is_dangerous_command (#33455)
## Summary

Cherry-picks the seven commits from
[openai/codex-internal#1942](https://github.com/openai/codex-internal/pull/1942)
onto `release/0.144`.

This backport:

1. enables dangerous-command detection in danger-full-access mode
2. expands literal Bash parsing so additional forced `rm` forms are
detected
3. returns a specific rejection reason to the model when a dangerous
command is denied

The commits applied without conflicts or release-only changes. `git
range-diff` confirms every cherry-picked patch matches the source PR
exactly.

## Validation

- `just test -p codex-shell-command` (141 passed)
- `just test -p codex-core exec_policy` (107 passed)
- `just test -p codex-core` outside the sandbox (2,947 passed; 4
unrelated environment/setup failures)
- 3 RMCP tests could not locate the test-only `test_stdio_server` binary
- 1 user-shell environment test observed the tool runner's mandated
`CODEX_SANDBOX_NETWORK_DISABLED=1`
- `just fix -p codex-core`
- `just fix -p codex-shell-command`
- `just fmt`
- `git diff --check`
2026-07-15 18:09:11 -07:00
rhan-oai
d82b7e5d4c Use model catalog policies for Guardian auto review (#32875)
- Add an `auto_review.policy` field to model catalog messages.
- Use the selected Guardian model's catalog policy for review-session instructions, while preserving the precedence of `guardian_policy_config` and falling back to the built-in policy when neither is present.
- Preserve auto-review messages when model instruction overrides remove catalog instruction templates.

- Cover configured-policy precedence, explicit empty catalog policies, catalog-message preservation, and propagation of the catalog policy into a prewarmed Guardian session.

GitOrigin-RevId: 26b61ae2958ea8325a64834dcf91f47e140d74b3
2026-07-13 21:14:18 -07:00
Felipe Coury
8a4d35a1e1 feat(tui): add an advanced reasoning picker 2026-07-13 02:11:55 -03:00
Dylan Hurd
32649bc5e6 [release/0.144] Revert "Update auto review prompting" (#32672)
## Summary

- Revert c27a909508e09105c80cc162e250e623a96a8f82 ("Update auto review
prompting") in full on the `release/0.144` branch.
- Restore the prior Guardian policy template, review request layout, and
tool specifications.
- Restore the corresponding Guardian tests and snapshots.

## Why

The auto-review prompting update needs to be rolled back from the 0.144
release line. This PR contains the direct, conflict-free Git cherry-pick
with no additional product changes.

## Validation

- `just fmt`
- `just test -p codex-core guardian::tests` (58 passed)
- `just test -p codex-core` (2,806 passed; 137 environment-sensitive
failures caused by sandbox restrictions on local ports and process
operations)
- `git show --check`
2026-07-12 20:13:56 -07:00
Channing Conger
77c42a202a code-mode: fall back to using in process v8 if we fail to resolve external process (#31899)
## Why

Not every Codex distribution currently includes the
`codex-code-mode-host` companion binary. Enabling the process-host
feature should not make code mode unavailable on those surfaces while
packaging support is being completed.

## What changed

- Fall back to an in-process code-mode session only when spawning the
companion binary returns `io::ErrorKind::NotFound`.
- Keep permission, handshake, timeout, and other host failures visible
instead of silently falling back.
- Store the provider's owned-process/in-process choice as one
enum-backed state so later sessions reuse the fallback decision.
- Preserve the underlying spawn `io::Error` while retaining the host
path in the displayed error.
- Update provider, `CodeModeService`, and end-to-end coverage to verify
successful fallback execution.

## Test plan

- `just test -p codex-code-mode`
- `just test -p codex-core missing_process_host`
2026-07-09 14:53:33 -07:00
Channing Conger
9d47eb221d code-mode: fix installation on darwin (#31876)
Installer needs to symlink code-mode-host next to codex on install.
2026-07-09 14:48:44 -07:00
efrazer-oai
7ef1728763 fix: parse compact release metadata in installer (#31667)
# Summary

GitHub's latest-release endpoint can return compact, single-line JSON.
The standalone installer treated release metadata as line-oriented text,
so those responses could make asset lookup fail even though the
requested assets were present.

The regression was introduced by
[#31056](https://github.com/openai/codex/pull/31056). That change reused
the `/releases/latest` metadata response for both version resolution and
asset lookup, exposing the existing formatting-sensitive asset parser to
compact responses from that endpoint.

This change parses the release metadata once with a one-pass POSIX awk
scanner. The scanner tracks JSON strings and nesting, extracts the root
release tag plus direct asset name/digest pairs, and produces the same
result regardless of whitespace or object field order. It uses POSIX
`fold` to bound awk record sizes so compact responses stay fast across
awk implementations.

Fixes #31520.

## Changes

- replace line-oriented release metadata matching with structure-aware
parsing
- reuse the parsed metadata for latest-version and asset-digest lookup
- add regression coverage for compact JSON, reordered fields, nested
decoys, and JSON-looking release text

## Design decisions

- Keep the installer dependency-free by using standard POSIX tools
already required by the shell installer.
- Parse only the GitHub release fields the installer consumes, in one
pass, instead of vendoring a general JSON library.
- Preserve asset-object boundaries so nested or string-encoded `name`
and `digest` fields cannot be mistaken for release assets.

## Testing

- Tests: focused installer suite locally and on Linux.
- Smoke tests: real pretty and compact GitHub release metadata,
latest-release resolution, and checksum-asset selection.
- Portability: macOS awk plus Linux gawk, mawk, and nawk.
- Stress coverage: randomized formatting and field order, adversarial
nested/string content, and a synthetic 2,000-asset compact response.
2026-07-09 14:48:39 -07:00
github-actions[bot]
3380969a29 Update models.json (#31684)
Automated update of models.json.

---------

Co-authored-by: sayan-oai <244841968+sayan-oai@users.noreply.github.com>
Co-authored-by: sayan-oai <sayan@openai.com>
2026-07-09 03:40:08 +00:00
Won Park
a7c72aee8b Use the image generation extension by default (#31596) 2026-07-09 12:25:19 +09:00
Owen Lin
2342b2c2a6 feat(rollout): persist TurnItems for paginated thread rollouts (#30188)
## Description

This PR makes new threads with `history_mode = "paginated"` persist
`ItemCompleted(item: <turn_item>)` in their rollout JSONL file.

Legacy threads keep persisting the existing legacy events. Because the
format is selected per thread, a rollout is either legacy or paginated;
we do not need to support mixed rollouts containing both
representations.

This PR depends on [#31473](https://github.com/openai/codex/pull/31473).

## Why

Paginated thread history needs stable turn/item IDs and completed item
snapshots so the later SQLite projector can materialize appended rollout
JSONL without rebuilding the whole thread.

Keeping the legacy persistence policy unchanged avoids changing
historical rollouts or the readers that still consume them.

## What changed

- Made rollout filtering history-mode aware. Paginated threads keep
completed canonical `ItemCompleted` events and drop their redundant
legacy projections; legacy threads keep the existing event set.
- Made forks inherit the source thread history mode, so copied legacy
history is never filtered as paginated.
- Made paginated threads assign IDs to locally-created response items
even when `Feature::ItemIds` is off, and reject streamed output items
that arrive without server IDs.
- Updated legacy turn replay, rollout list/search, and SQLite metadata
extraction to understand completed canonical user-message items.
2026-07-08 19:55:03 -07:00
Michael Bolin
5892c7b69d model-provider: route model discovery through HTTP client factory (#31361) 2026-07-09 02:32:10 +00:00
stevenlee-oai
555aa79d5a [connectors] Refresh codex_apps /ps/mcp auth (#31486)
[Codex Thread
019f2408-dc59-79f2-b245-4c11debd1a61](https://codex-thread-link.openai.chatgpt-team.site/thread/019f2408-dc59-79f2-b245-4c11debd1a61)

## Why

Long-lived Codex sessions can outlive the ChatGPT bearer token that was
present when the MCP runtime started.

The Responses path already recovers from token expiration by refreshing
or reloading the shared `AuthManager`. The reserved `codex_apps`
hosted-plugin client did not observe that update: `McpConnectionManager`
built its `/ps/mcp` HTTP auth once from a `CodexAuth` snapshot, and
`auth_provider_from_auth` copied that snapshot bearer into a static
`BearerAuthProvider`.

After the copied bearer expired, `/ps/mcp` kept sending it even though
Responses had a newer token in the same `AuthManager`. The failure
occurred before downstream connector execution, so unrelated apps such
as Gmail, Slack, and Google Calendar could all fail with the same
transport-level `401 token_expired`.

This replaces
[openai/codex#29474](https://github.com/openai/codex/pull/29474), which
was closed for inactivity without being merged. A new long-lived-session
report reproduced the same simultaneous `/ps/mcp` expiry pattern across
unrelated apps.

## What changed

- Add an `AuthManager`-backed request-header provider in
`codex-model-provider`. It keeps an `Arc<AuthManager>` and reads
`auth_cached()` for each outbound request, so the next `/ps/mcp` call
sees a token refreshed by the existing Responses/auth-recovery flow.
- Scope that provider to the startup account, ChatGPT user, and
workspace identity. Same-identity token reloads are followed; an account
switch emits no ambient auth until account-scoped MCP state is rebuilt.
- Have `McpConnectionManager` construct the dynamic provider only for
the reserved `codex_apps` registration used by the hosted-plugin
`/ps/mcp` path.

| MCP path | Auth behavior after this change |
| --- | --- |
| Reserved `codex_apps` hosted-plugin `/ps/mcp` | Read current
same-identity auth from the shared `AuthManager` per request |
| `codex_apps` with `CODEX_CONNECTORS_TOKEN` | Keep the environment
bearer-token override |
| User-configured/direct MCP registrations | Keep their existing
configured auth path |

## Non-goals

- No plugin-service changes.
- No downstream Slack, Gmail, Calendar, or other connector
OAuth/link-refresh changes.
- No auth UI changes.
- No behavior change for user-configured/direct MCP registrations.
- No new `/ps/mcp`-initiated token refresh; this makes `/ps/mcp` observe
refreshes already performed through the shared `AuthManager`.

## Tests

- `just test -p codex-model-provider`
- Covers same-identity token reloads and refuses a changed startup
identity.
- `just test -p codex-mcp`
- `just test -p codex-core mcp_auth_refresh`
- Creates the reserved hosted-plugin `codex_apps` `/ps/mcp` client
before the shared `AuthManager` changes, updates that same manager
through its public external-auth path, performs a real `tools/call`, and
asserts the request uses the current bearer.
2026-07-08 21:46:11 -04:00
pakrym-oai
20e5edfa74 Expand agent core ownership (#31675)
The agent core team owns the core agent implementation and should review
changes to its adjacent runtime crates. Those crates were not covered by
the existing CODEOWNERS rules.

This adds `@openai/codex-core-agent-team` ownership for:

- `codex-rs/arg0`
- `codex-rs/codex-mcp`
- `codex-rs/exec-server`

## Validation

- `git diff --check`

No runtime tests are applicable because this only changes CODEOWNERS
metadata.
2026-07-08 18:04:44 -07:00
viyatb-oai
0746e8a345 [codex] Preserve reviewer when resuming threads (#30278)
## Why

A thread resumed without an explicit reviewer could pick up the reviewer
from the current config instead of preserving the reviewer already in
use by the thread. After an app restart, this meant a thread running
with auto review could silently switch back to user review, and the next
turn could continue under the wrong reviewer.

## What changed

Persist the effective reviewer with each turn and restore the latest
persisted value when the thread resumes. If the resume request
explicitly provides a reviewer, that value still takes precedence.

## Test plan

- Added a regression test that starts a thread with auto review, records
a turn, restarts with user review in config, resumes without an
override, and verifies that auto review is preserved.
- `just test -p codex-protocol`
- `just test -p codex-state`
- `just test -p codex-rollout`
- `just test -p codex-app-server
thread_resume_preserves_persisted_approvals_reviewer`
- Clippy for the affected crates
2026-07-09 00:58:28 +00:00
Adam Perry @ OpenAI
3fa90665fe test: add delayed exec-server transport (#31427)
## Why

Macrobenchmarks benefit from having a way to exercise remote-executor
latency without depending on Docker.

This is a very minimal first cut, if we find that simulating network
conditions is useful we can always expand this scope or switch to a more
robust network shaping approach.

## What

- add a package-local exec-server binary for Cargo and Bazel test
fixtures
- add a host-local WebSocket exec-server fixture and fixed-delay
interposer
- let TestAppServer route its auto environment through that delayed
WebSocket transport
- cover the delayed thread/start path through the public app-server API

## Stack

1. [#31425 test: add TestAppServer
builder](https://github.com/openai/codex/pull/31425)
2. [#31427 test: add delayed exec-server
transport](https://github.com/openai/codex/pull/31427)
3. [#31295 bench: add cold skill load
macrobenchmark](https://github.com/openai/codex/pull/31295)
4. [#31428 bench: add e2e benchmark
entrypoints](https://github.com/openai/codex/pull/31428)
5. [#31429 ci: smoke Bazel e2e
benchmarks](https://github.com/openai/codex/pull/31429)
2026-07-09 00:17:38 +00:00
github-actions[bot]
b780738014 Update models.json (#21818)
Automated update of models.json.

---------

Co-authored-by: aibrahim-oai <219906144+aibrahim-oai@users.noreply.github.com>
Co-authored-by: Ahmed Ibrahim <aibrahim@openai.com>
Co-authored-by: Sayan Sisodiya <sayan@openai.com>
2026-07-08 16:24:12 -07:00
olliem-oai
3eb56537eb Update auto review prompting (#31480)
## Why
Auto-review performance is weaker because of confusing instructions
about sandbox permissions, and because it is given many tools which are
irrelevant to it.

## What
* Update the auto review prompt
* Remove the permissions_instructions developer message
* Only pass exec_tool and view_image tool to the reviewer

## Validation
`just fmt`
`cargo test -p codex-core --lib --quiet`
2026-07-08 23:15:43 +00:00
jacobzhou-oai
a09a7c41d8 [codex-apps] Omit internal fields from file payloads (#31330)
## Summary

Codex Apps file parameters are exposed to the model as local paths,
uploaded at execution time, and rewritten into provided-file payloads
before the MCP tool call.

The rewrite currently forwards two internal upload fields, `uri` and
`file_size_bytes`, even though they are not part of the documented app
file-reference shape. Strict app schemas can reject those extra fields
before execution.

## Changes

- Stop copying `uri` and `file_size_bytes` into app-facing MCP
arguments.
- Keep the internal `UploadedOpenAiFile` result unchanged.
- Preserve the existing `download_url`, `file_id`, `mime_type`, and
`file_name` behavior for scalar and array file inputs.
- Verify the MCP invocation and post-tool hook receive exactly the
documented four-field payload against an `additionalProperties: false`
schema.

This intentionally does not add schema inspection or change how
`openai/fileParams` names are discovered.

## Validation

- `just test -p codex-core mcp_openai_file` (6 passed)
- `just test -p codex-core codex_apps_file_params_` (2 passed)
- `just fix -p codex-core`
- `just fmt`
- `git diff --check`
2026-07-08 16:04:30 -07:00
Felipe Coury
4e270ddec4 test(app-server): use native rollout fixture paths (#31663)
## Why

Windows CI now places temporary and build files on the `D:` Dev Drive.
Fake rollout metadata still stored `/` as its working directory, but `/`
is drive-relative on Windows. When the migrated auto-environment tests
resumed or listed those rollouts, the fixture resolved to `D:\` while
the established test expectation remained `C:\`, causing unrelated PRs
to fail the Windows app-server shard.

This follows the interaction between #31357, which moved CI build paths
to the Dev Drive, and #31614, which migrated these app-server tests to
automatic environments.

## What

- Construct fake rollout working directories with `test_path_buf("/")`,
producing a fully qualified native path on Windows while preserving `/`
on Unix.
- Use the same native test-path helper for the legacy
conversation-summary expectation.

## How to Test

Automated tests were intentionally not run locally at request; the
app-server suite was stopped during compilation. `just fmt` completed
successfully.

To verify the regression on a Windows runner:

1. Configure `TEMP` and `TMP` on a non-`C:` drive, as CI does with the
Dev Drive.
2. Run `just test -p codex-app-server`.
3. Confirm the existing thread list, read, and resume tests no longer
report `D:\` actual versus `C:\` expected paths.

This is a test-fixture-only change, so there is no product smoke path.
2026-07-08 19:55:05 -03:00
Eric Traut
f3bfaca3c1 Clarify device-code phishing warning (#31648)
## Why

The existing device-code warning does not help users distinguish a login
they initiated from a phishing attempt. The warning should tell users to
stop when the code came from a website or another person.

## What changed

- Updated the warning in the direct CLI and TUI device-code login flows
with actionable guidance.
- Added focused coverage for the styled direct CLI prompt.
2026-07-08 15:45:57 -07:00
Channing Conger
c55cb4b363 code-mode: make all approvals trigger elicitation pause (#31650)
### summary

We want to pause code-mode from yielding back to the model when a
subcommand triggers an approval prompt. This means that all of these
previously inline blocking requests should also take out a
ElicitationService registration.

This also does some plumbing refactoring to request patch approval to
make it match the other `request_*_approval` methods in that it blocks
on the approval in the function instead of returning the oneshot
channel, this affords our ability to encapsulate the ElicitationService
registration via RAII.

Adds tests to confirm the blocking behavior for code_mode both in suite
tests and that the session holds them.
2026-07-08 15:27:04 -07:00
fbauer33
b6f9aee16d [codex] increase tool schema compaction threshold (#31497)
## Why

The 4,000-byte limit is compacting the tool schemas of some hero
usecases.

## What changed

Raise the limit to 5,000 bytes and update compaction test fixtures
accordingly.
2026-07-08 15:14:23 -07:00
Adam Perry @ OpenAI
5c7624a69e test: migrate app-server v2 starts to auto env (#31614)
## Why

We should be running as many integrations tests as possible against the
split cross-OS configuration.

## What

- migrate eligible thread starts and builders to auto env
- keep explicit custom/local-executor cases local with rationale
comments
- keep auto-env coverage where possible and add narrow `TODO(anp)` skips
for fixtures that are not target-native yet
2026-07-08 14:13:59 -07:00
Adam Perry @ OpenAI
4b64bf0751 chore: remove inert cargo audit workflow (#31461)
## Why

`codex-rs/.github/workflows/cargo-audit.yml` is nested below the
repository root, so GitHub Actions never discovers or runs it. RustSec
advisory enforcement already runs through the root `cargo-deny`
workflow.

## What

Remove the inert nested Cargo audit workflow.

## Validation

- Ran `git diff --check`.
- Verified the workflow is absent from GitHub registered workflows and
that root blocking CI invokes `cargo-deny`.
2026-07-08 14:07:43 -07:00
Adam Perry @ OpenAI
bd5c860abe ci: route build IO through Dev Drives (#31357)
## Why

Windows Cargo and Bazel jobs spend significant time in filesystem-heavy
build and cache directories. Route those directories through one CI
build root so Windows can use its Dev Drive and Unix can use a stable
cache root.

## What

- Have `setup-ci` define `CI_BUILD_ROOT`, `CARGO_TARGET_DIR`, Bazel
cache/output paths, and temp paths.
- Require Windows to find or provision a verified Dev Drive instead of
falling back to `C:`.
- Pass the shared Bazel output base to `setup-bazel` so its explicit
`output_base` does not defeat Dev Drive routing.
- Point nextest, release, and V8 source-build paths at the shared
environment contract.

## Benchmark results

One-off cold-cache WPR/ETW traces show the explicit Bazel output-base
routing removes the dominant `C:` traffic:

| sample | `C:\_bazel` | summed `C:` traffic | traced test step |
|---|---:|---:|---:|
| shard 1 before | 62.2 GiB | 85.2 GiB | 16m22s |
| shard 1 updated | 0 | 16.5 GiB | 12m05s |
| shard 3 before | 67.2 GiB | 84.6 GiB | 16m48s |
| shard 3 updated | 0 | 13.5 GiB | 11m08s |

For a cold x64 V8 source build, the retained build-tail sample showed
`D:\cargo-target` at ~1.29 GiB while measured `C:` roots totaled ~0.45
GiB (`C:\Users` ~0.33 GiB, `C:\Program Files` ~0.06 GiB, `C:\Windows`
~0.03 GiB). The full cold build took 2h20m36s.

The Bazel timing improvement is directional because both refreshed
shards failed tests. The V8 trace is a bounded build-tail sample, not
the full build. All final samples had zero lost ETW events; VHDX traffic
was excluded from the optimization ranking.

Runs: [baseline
Bazel](https://github.com/openai/codex/actions/runs/28911908527),
[updated
Bazel](https://github.com/openai/codex/actions/runs/28917133701), [V8
build tail](https://github.com/openai/codex/actions/runs/28933626678).

## Manual validation

- Ran `just fmt`.
- Ran `just test-github-scripts` (35 tests).
- Parsed GitHub Actions YAML with `yq`.
- Ran `git diff --check`.

## Stack

- [#31332](https://github.com/openai/codex/pull/31332) — parameterize
Cargo target paths
- [#31356](https://github.com/openai/codex/pull/31356) — Windows 2025
runner bump
- [#31357](https://github.com/openai/codex/pull/31357) — Dev Drive I/O
routing
2026-07-08 14:06:37 -07:00
Michael Bolin
e621d7df8c core: preserve Responses WebSockets with system proxy (#31441)
## Why

Responses WebSockets are the normal lower-latency transport for
WebSocket-capable providers. They must not bypass an OS-selected proxy
when `features.respect_system_proxy` is enabled, but disabling
WebSockets whenever the feature is enabled would impose a substantial
performance penalty.

Merged PR #31622 introduced the reusable proxy-aware WebSocket
transport. This PR makes the Responses API its first consumer so the
existing fast path uses the same effective proxy and trust policy as
HTTP.

## What changed

- Register `codex-websocket-client` as a workspace dependency and use it
from `codex-api`.
- Feed the shared crate’s route-independent `WebSocketConnection` into
the existing Responses message pump.
- Require a configured `HttpClientFactory` for normal Responses
WebSocket connections and the CLI doctor probe, so neither path can open
a connection without consulting the effective proxy policy.
- Pass the session factory from `core` and the effective configuration
factory from `doctor`.
- Add an end-to-end Responses test that enables `RespectSystemProxy`,
asserts the resolved policy, completes a turn over WebSocket, and
verifies the connection and request counts.
- Keep the existing Responses protocol handling, ping/pong pump, and
session-scoped HTTP fallback unchanged.

The DNS, proxy, TLS, custom-CA, and Happy Eyeballs implementation and
its transport tests live in merged PR #31622. This PR deliberately
contains only the Responses integration and does not duplicate that
transport code.

## Review guide

1. `codex-rs/codex-api/src/endpoint/responses_websocket.rs` constructs
the shared connector and adapts its uniform stream to the existing pump.
2. `codex-rs/core/src/client.rs` supplies the session-scoped factory for
production Responses connections.
3. `codex-rs/cli/src/doctor.rs` supplies the effective configuration
factory to the handshake probe.
4. `codex-rs/core/tests/suite/client_websockets.rs` covers the
enabled-feature path end to end.

## Test plan

- `cargo check --tests -p codex-api -p codex-core -p codex-cli`
- `just test -p codex-api`
- `just test -p codex-core
responses_websocket_streams_with_system_proxy_feature`
- `cargo shear`
- `just bazel-lock-check`


---
[//]: # (BEGIN SAPLING FOOTER)
Stack created with [Sapling](https://sapling-scm.com). Best reviewed
with [ReviewStack](https://reviewstack.dev/openai/codex/pull/31441).
* #31637
* #31431
* #31363
* #31362
* #31361
* __->__ #31441
2026-07-08 14:06:15 -07:00
Owen Lin
602dbb42dc core: stop emitting legacy command events directly (#31629)
## Description

This PR removes the last path in core that emits `ExecCommandBegin` /
`ExecCommandEnd` directly.

Every command execution now starts and completes through canonical
`ItemStarted` / `ItemCompleted(TurnItem::CommandExecution)`. The
existing `HasLegacyEvent` compatibility layer still fans out Begin/End
afterward, so raw core event consumers and legacy rollout replay keep
seeing the same events.

`UnifiedExecInteraction` is dormant today. Live unified exec uses
`UnifiedExecStartup` for command lifecycle and `TerminalInteraction` for
`write_stdin` and polling, so this is code cleanup rather than a current
product behavior change. The main win is the code-level invariant where
all core flows emit `TurnItem` instead of legacy events.

## What changed

- Removed the `UnifiedExecInteraction` branches that emitted legacy
command events directly.
- Routed every command source through the existing canonical
`CommandExecution` lifecycle and compatibility fanout.
2026-07-08 20:53:36 +00:00
Celia Chen
bff9c4945a feat: change amazon Bedrock GPT-5.6 display names (#31636)
## Why

Amazon Bedrock's GPT-5.6 variants currently appear as only `Sol`,
`Terra`, and `Luna`. Those labels omit the model family and version,
making them ambiguous in model lists and inconsistent with the naming of
other GPT models.

## What changed

- Rename the three Bedrock model display names to `GPT-5.6 Sol`,
`GPT-5.6 Terra`, and `GPT-5.6 Luna`.
- Strengthen the Bedrock model-manager test to verify both model IDs and
their propagated display names.

Model IDs, ordering, priorities, reasoning support, and default
selection are unchanged.
2026-07-08 13:53:18 -07:00
Owen Lin
ba5dd1fd3a feat(core): emit canonical hook prompt items (#31630)
## Description

This PR moves hook prompts onto the canonical `TurnItem` lifecycle in
core.

Stop hooks now record their `ResponseItem` through the existing
lifecycle path, which emits `ItemStarted` and `ItemCompleted`.
App-server consumes those events directly instead of deriving a hook
prompt from `RawResponseItem`.

## Why

Hook prompts were the only `ThreadItem` app-server synthesized from
`RawResponseItem`. This brings them in line with other core-owned turn
items while preserving legacy rollout replay.

## What changed

- Route stop-hook prompts through
`record_response_item_and_emit_turn_item`.
- Materialize canonical hook prompts in `ThreadHistoryBuilder`.
- Remove `RawResponseItem` to `ThreadItem` synthesis while preserving
legacy rollout replay.
- Add focused coverage for lifecycle emission and canonical and legacy
history materialization.
2026-07-08 13:42:06 -07:00
sayan-oai
bdaad6820c Reuse MCP tool snapshot within a sampling request (#31292)
Follow-up to #30226.

## Why

#30226 makes Apps World State inspect the MCP tool list, while
tool-router construction reads the same list again later in the sampling
request. `list_all_tools()` walks the MCP clients and may reconnect or
wait for tools, so doing that work twice adds latency and lets context
and tool construction observe different MCP states for one request.

## What

- Add a lazy MCP tool snapshot to `StepContext`.
- Reuse that snapshot for Apps World State and tool-router construction.
- Let each new `StepContext` refresh naturally for the next sampling
request, without manager-level caching or invalidation.

## Testing

- `just test -p codex-core apps_instructions`
- `just test -p codex-core
apps_guidance_appears_after_background_recovery_within_a_turn`
2026-07-08 13:35:31 -07:00
Abhinav
c6b124bb31 [codex] Grant Windows sandbox access to primary runtime (#31574)
## Why

Codex Desktop installs its managed primary runtime under
`%USERPROFILE%\.cache\codex-runtimes`. Elevated Windows sandbox commands
run as dedicated sandbox users. The synchronous runtime ACL refresh
repairs read/execute access for the Desktop runtime directories under
`%LOCALAPPDATA%\OpenAI\Codex`, but did not include the managed primary
runtime cache.

As a result, the Desktop app could discover a bundled runtime while a
sandboxed command received `ACCESS_DENIED` when reading or executing it.

## What changed

- Include `%USERPROFILE%\.cache\codex-runtimes` in the managed runtime
paths considered by the Windows sandbox ACL refresh.
- Reuse the existing inherited read/execute ACL repair; no write
permission is added.
- Add Windows-target regression coverage for the runtime path list and
the primary-runtime-only case.

## Impact

Bundled Python, Node, and native tools remain usable from elevated
Windows sandbox sessions without broadening write access or granting
access to the rest of the user profile.

## Validation

- `just fmt`
- `just test -p codex-windows-sandbox` (10/10 host-side tests passed)
- Windows-target path tests included for CI
2026-07-08 13:22:10 -07:00
Michael Bolin
2780bd588f websocket-client: add proxy-aware connector (#31622)
## Why

The route-aware WebSocket connection setup in #31441 is transport
infrastructure rather than Responses API protocol logic. Landing it
first in a dedicated crate keeps `codex-api` focused on request and
response behavior and makes the transport reusable by future WebSocket
clients.

WebSockets must also apply the same effective outbound proxy and
custom-CA policy as HTTP without disabling the lower-latency WebSocket
path. Requiring an `HttpClientFactory` when constructing the connector
makes proxy-policy resolution part of the API instead of an optional
call-site convention.

This PR is an independent prerequisite based directly on `main`. After
it merges, #31441 can rebase onto it and replace its in-crate connector
with this API.

## What changed

- Add a new `codex-websocket-client` workspace crate with a
`WebSocketConnector` constructed from the effective `HttpClientFactory`.
- Resolve every destination through that factory before connecting, then
support direct connections, transport-default routing, HTTP proxies, and
TLS-encrypted HTTPS proxies.
- Preserve custom-CA trust for proxy and target TLS handshakes and
preserve Happy Eyeballs fallback for explicit direct and proxy routes.
- Expose an established `WebSocketConnection` as a uniform `Stream` and
`Sink`, hiding route-specific transport types from protocol clients.
- Add focused integration-style coverage for the public connector and
message stream, real WSS over direct and CONNECT routes, implicit and
explicit HTTPS proxy ports, and stalled-address-family fallback.

## Review guide

1. `codex-rs/websocket-client/src/lib.rs` defines the small public API
and the factory-required policy invariant.
2. `codex-rs/websocket-client/src/dialer.rs` contains DNS, TCP, proxy
tunneling, TLS, and WebSocket handshake setup.
3. `codex-rs/websocket-client/src/dialer_tests.rs` verifies the public
stream, direct and proxied WSS paths, HTTPS port preservation, and Happy
Eyeballs timing.
4. There is intentionally no consumer migration here; #31441 will become
the first consumer after this prerequisite merges.

## Test plan

- `cargo check -p codex-websocket-client --tests`
- `just test -p codex-websocket-client`
- `cargo shear`
- `just bazel-lock-check`
2026-07-08 12:57:27 -07:00
Shijie Rao
927004c06d tui: warn on Ultra with high multi-agent concurrency (#31621)
## Why

Ultra reasoning may proactively use multiple agents. When
`features.multi_agent_v2.max_concurrent_threads_per_session` is
configured at 8 or higher, explicitly selecting Ultra can allow up to `N
- 1` subagents to work concurrently and increase usage quickly. Showing
the configured limits at selection time makes that tradeoff visible and
points users to the setting that controls it.

## What changed

- Show a warning history cell after the user explicitly selects Ultra
reasoning with a concurrent-thread limit of at least 8.
- Include the configured concurrent-thread count and maximum subagent
count in the warning.
- Apply the warning consistently across model and reasoning pickers,
Plan-mode scope selection, and reasoning shortcuts.
- Keep the trigger limited to the selected reasoning effort and
configured thread limit, independent of how multi-agent v2 is activated.
- Add focused threshold coverage and an `insta` snapshot for the
rendered warning.

## User impact

Users with high multi-agent concurrency receive a concrete warning
immediately after selecting Ultra. Other reasoning efforts, limits below
8, and startup behavior are unchanged.


https://github.com/user-attachments/assets/5795bcb9-432e-42bc-bde3-ee363e9aeb69

## Test plan

- `just test -p codex-tui ultra_reasoning_selection`
- Built the full debug CLI with `cargo build -p codex-cli`.
- Manually verified the debug TUI with a 10-thread limit: switching from
Max to Ultra displayed the warning with 10 concurrent threads and up to
9 subagents.
2026-07-08 12:29:21 -07:00
stefanstokic-oai
1d14221af2 [codex] Sanitize imported session fallback titles (#29875)
## Summary

- sanitize Claude-derived fallback titles without changing explicit
custom or generated source titles
- recognize known leading control wrappers, including `ide_opened_file`,
`ide_selection`, `local-command-stdout`, and `local-command-stderr`
- skip separate control-only user records and use the next meaningful
user message as the fallback title
- use `Imported session` only when every user record is control-only
- preserve raw imported messages, previews, and transcript provenance
unchanged

## Why

Sessions without an explicit source title fall back to user-message
content. Claude metadata can arrive as separate leading user records, so
stripping only the first record still exposed raw control markup or
produced an unhelpful fallback.

## Impact

Imported chats now receive readable fallback titles such as `Fix auth
flow`, while the original wrapper-bearing messages remain visible in the
imported transcript.

## Validation

- `just test -p codex-external-agent-sessions` — 37 tests passed
- `just test -p codex-app-server
external_agent_config_import_creates_session_rollouts` — focused
integration test passed
- `just fix -p codex-external-agent-sessions`
- `just fix -p codex-app-server`
- `just bazel-lock-check`
- `just fmt`
2026-07-08 15:20:40 -04:00
jif
a14b4c2d7f Bound exec-server pending RPCs (#31578)
## Why

An untrusted exec-server can stop reading requests, never answer them,
or send guessed responses before queued requests are written. Without
client-side admission, the orchestrator can retain unbounded RPC call
futures and request payloads.

We need a hard bound without adding a blanket timeout, because
individual operations already own their timeout and cleanup semantics.

## What changed

- hold one of 64 shared admission permits for the full lifetime of each
regular RPC call, so an early response cannot free capacity while the
call remains queued
- allow `process/terminate` and `fs/close` to use one additional cleanup
permit
- close the transport and fail pending calls if the cleanup permit is
also stuck, allowing teardown or recovery to release remote resources
- leave the wire format and existing timeout behavior unchanged

## Tests

- `rpc_client_call_has_no_implicit_deadline` verifies that ordinary
calls remain untimed
- `rpc_client_bounds_in_flight_calls_and_preserves_cleanup` covers the
regular-call cap, guessed early responses, cleanup admission, and the
cleanup circuit breaker
2026-07-08 19:46:18 +01:00
Francis Chalissery
8d83687aaa Fall back to HTTP when Apple Git is unavailable (#31496)
## Summary

- on macOS, resolve the curated plugin sync Git executable without
executing it
- treat Apple’s `/usr/bin/git` shim as unavailable when `xcode-select
-p` reports that developer tools are absent
- skip directly to the existing GitHub HTTP fallback in that case
- preserve the original `git` command lookup on Windows and Linux,
including CCA

## Root cause

Curated plugin startup sync invokes `git ls-remote` before its HTTP
fallback. On a clean Mac, `git` resolves to Apple’s `/usr/bin/git` shim,
and executing the shim opens the Xcode Command Line Tools installer
before the process can fail and reach HTTP.

On macOS, this change resolves Git through `PATH` without executing it.
If the selected binary is Apple’s shim and developer tools are
unavailable, startup sync marks the Git transport unavailable and enters
the existing HTTP fallback immediately.

The new availability detection is macOS-only by construction. Windows
and Linux still execute the literal `git` command as before. If Git is
missing on Windows, the existing spawn-error path falls back to HTTP;
Linux/CCA receives no new lookup or startup behavior.

## Eager Git audit

I also audited production Git process spawns in `codex-rs`.

- This curated catalog sync is the only default projectless app-server
startup path found.
- Configured Git marketplace auto-upgrade runs Git at plugin startup,
but only after a user has explicitly configured a Git marketplace.
- Experimental Memories has background Git metadata/baseline paths when
the feature is enabled.
- The separate cloud-tasks UI probes Git during environment
autodetection.
- Normal thread/turn Git metadata is gated by filesystem discovery of an
existing `.git` entry.
- Marketplace add/install, patch apply, doctor, and TUI `/diff` paths
are explicitly user-invoked.

## Validation

- `just fmt`
- `just bazel-lock-update` — succeeded with no lockfile delta
- `just test -p codex-core-plugins` — 313 passed
- `just fix -p codex-core-plugins` — completed; emitted one pre-existing
unrelated `large_enum_variant` warning in `manifest.rs`
- `git diff --check`
2026-07-08 18:38:18 +00:00
Owen Lin
154433c8f0 chore(protocol): use UUIDv7 for generated item IDs (#31524)
## Description

Use UUIDv7 for item IDs generated locally in `codex-protocol`.

This covers user messages, agent messages, hook prompts, and context
compaction items. These IDs are generated once and carried through item
lifecycle, so this only changes the UUID version. It keeps generated
item IDs consistent with thread and turn IDs.
2026-07-08 18:35:42 +00:00
Eric Traut
9077300d7f tui: sanitize terminal controls in user messages (#31494)
## Why

A pasted user message can contain a raw CSI sequence that corrupts
terminal scrollback, making the latest response or input appear missing.
The same sequence may already be present in persisted history when an
affected conversation is resumed.

## What changed

- Remove CSI sequences and non-whitespace control characters from
explicit paste input before it enters the composer.
- Apply the same small transform before rich and raw `UserHistoryCell`
rendering, so existing conversations render safely without mutating
stored message content.
2026-07-08 11:10:35 -07:00
Channing Conger
9c6715924b code-mode: move to hosted mode by default (#31500)
## Summary

  - Promote code_mode_host to stable and enable it by default.
- Preserve features.code_mode_host = false as an opt-out to the
in-process runtime.
  - Run core code-mode tests through the standalone host.
  - Keep explicit coverage for missing-host failures.
2026-07-08 11:06:58 -07:00
Felipe Coury
166534fc22 fix(windows-sandbox): allow deletion in writable roots (#31138)
## Why

The legacy unelevated Windows sandbox allowed tools to create and update
files in workspace-write roots, but it could not delete files that
already existed there. This breaks operations such as `apply_patch` file
deletion and replacement in the workspace, `TEMP`, and `TMP`.

The delete grant must also preserve deny-write carveouts. Granting
`FILE_DELETE_CHILD` on a writable parent would let the sandbox remove
protected children such as `.git` or an explicit read-only subpath even
when those children have direct deny ACEs.

This addresses the delete-failure variant reported in #30009 and #30712.
It does not claim to fix their separate split-root setup,
elevated-helper, or proxy-related failure modes.

## What Changed

- Give writable-root capability ACEs inheritable `DELETE` rights without
granting parent-level `FILE_DELETE_CHILD`, so descendants can be removed
while protected children remain protected.
- Replace stale write ACEs that still contain `FILE_DELETE_CHILD`, and
make elevated setup detect and refresh that unsafe legacy state.
- Keep read-only capability handling unchanged.
- Add Windows regressions covering pre-existing files in the workspace,
`TEMP`, and `TMP`, plus protected `.git` and outside-root controls.

The core ACL behavior is in
[`acl.rs`](767540eec3/codex-rs/windows-sandbox-rs/src/acl.rs (L303-L438)),
stale-ACE detection is in
[`setup_main/win.rs`](767540eec3/codex-rs/windows-sandbox-rs/src/bin/setup_main/win.rs (L163-L179)),
and the end-to-end regression is in
[`unified_exec/tests.rs`](767540eec3/codex-rs/windows-sandbox-rs/src/unified_exec/tests.rs (L458-L568)).

## How to Test

On Windows:

1. Start Codex with the legacy unelevated Windows sandbox and a
workspace-write permission profile.
2. Seed pre-existing files in the workspace, `TEMP`, and `TMP`; also
create a sibling file outside the writable roots and a protected `.git`
directory.
3. Delete the three files inside writable roots through a sandboxed
command or `apply_patch`.
4. Confirm the writable-root files are deleted, while the outside-root
file and protected `.git` directory remain intact.

Targeted tests:

- `just test -p codex-windows-sandbox`
- Windows-only
`legacy_workspace_write_delete_is_limited_to_writable_roots`
- Windows-only `write_root_refresh_replaces_stale_delete_child_grant`

The final SHA passed all 31 required checks, including the Windows Bazel
test matrix, in [run
28886245161](https://github.com/openai/codex/actions/runs/28886245161).
2026-07-08 15:05:57 -03:00
jif
48cf582331 Round MCP timeout durations in error messages (#31612)
## Summary

MCP operation timeout errors currently print the full debug precision of
the remaining timeout budget. That makes a configured 30-second timeout
show up as something like `29.999999875s`.

This PR rounds the displayed duration to a whole unit, so the error is
clean and stable:

```text
timed out awaiting tools/list after 30s
```

The timeout behavior itself is unchanged; this only affects the
human-facing error text. A regression test covers the reported
`tools/list` case.
2026-07-08 18:58:02 +01:00
Adam Perry @ OpenAI
f73a072246 test: remove TestAppServer constructors (#31452)
## Why

Finish the TestAppServer builder migration after every caller has moved
off the compatibility constructors.

## What

- remove the obsolete public TestAppServer constructors
- leave TestAppServer::builder() as the only fixture construction API

## Validation

- cargo check -p codex-app-server --tests

## Cleanup stack

1. [#31425 test: add TestAppServer
builder](https://github.com/openai/codex/pull/31425)
2. [#31451 test: migrate TestAppServer callers to
builder](https://github.com/openai/codex/pull/31451)
3. this PR
2026-07-08 17:50:18 +00:00
Owen Lin
23aac925e7 feat(core): emit canonical review mode items (#31473)
## Description

This PR moves review-mode markers onto canonical `TurnItem` lifecycle:

- `TurnItem::EnteredReviewMode`
- `TurnItem::ExitedReviewMode`

Core now emits `ItemStarted` / `ItemCompleted` for both. The completed
items map back into the existing `EnteredReviewMode` /
`ExitedReviewMode` events, so raw core event consumers and legacy
rollout persistence keep seeing the old events.

This is the compatibility layer needed before paginated rollouts persist
review markers as `ItemCompleted(TurnItem)`.

## Why

Review markers were one of the remaining app-server thread items created
directly from legacy events. Giving them canonical items lets paginated
history persist stable turn/item IDs without changing legacy rollouts.

## What changed

- Added canonical review-mode `TurnItem`s and switched review flow to
emit their lifecycle.
- Added completed-item → legacy review event mappings with stable
turn/item IDs.
- Switched app-server live notifications to the generic canonical item
path and kept legacy replay compatible with old payloads.
- Updated `ThreadHistoryBuilder` to replay canonical review items even
though review turns still do not emit `TurnStarted`.
2026-07-08 09:59:50 -07:00
Owen Lin
a219b6fdb4 core: migrate standalone web search to extension-owned turn items (#31525)
## Description

This PR migrates standalone web search onto the extension-owned
turn-item path introduced in #31283.

Standalone web search now emits `ExtensionItem::WebSearch` through
generic `TurnItem::Extension`, while app-server still exposes the
existing typed `ThreadItem::WebSearch` JSON shape. Hosted Responses API
web search stays on core-owned `TurnItem::WebSearch`.

## What changed

- Added `web_search::WebSearchItem` and `WebSearchAction` to
`codex-extension-items` under the stable `web.search` kind.
- Collapsed `ExtensionTurnItem` to generic `{ item, legacy_events }` now
that no typed extension special cases remain.
- Kept the existing `WebSearchBegin` / `WebSearchEnd` compatibility
events and canonical-first ordering.
- Updated app-server projection/history and generated TypeScript; the
app-server JSON schema is unchanged.
2026-07-08 09:40:00 -07:00
Charlie Marsh
e212cc95b0 Detect Codex installs managed by pnpm (#31503)
## Why

The Codex JavaScript shim currently distinguishes npm and Bun installs,
but a global pnpm install falls back to npm. That causes the native CLI,
`codex doctor`, and update flows to report or run npm commands even
though pnpm owns the installation. pnpm also allows its global package
and bin directories to differ, so `PNPM_HOME` does not reliably identify
the package owner.

## What changed

- detect pnpm-managed installs by finding pnpm's
`node_modules/.modules.yaml` metadata from the launched JavaScript
entrypoint
- pass a mutually exclusive `CODEX_MANAGED_BY_PNPM` marker into the
native CLI
- represent pnpm in install context, doctor output, version checks, and
TUI update actions without changing the existing public
`InstallContext::from_exe` signature
- recommend `pnpm add -g @openai/codex` for pnpm-managed installs and
snapshot the rendered TUI update notice

Validated with the 9-test install-context suite, the focused pnpm TUI
snapshot test, and Node's syntax check for the launcher.

Closes https://github.com/openai/codex/issues/10294.
2026-07-08 12:17:58 -04:00
jif
0bbea86a6a Stabilize shared rollout budget test (#31587)
## Why

`subagent_usage_draws_from_the_shared_budget` intermittently fails even
when the shared-budget behavior is correct. `ResponseMock` records a
request before the custom Wiremock predicate is checked, so the
follow-up mock can also contain unrelated requests. In a [recent Windows
ARM64 run](https://github.com/openai/codex/actions/runs/28916285431),
`single_request()` saw all seven requests from the scenario.

## What changed

Select the request containing the follow-up user prompt before making
assertions. The test still requires exactly one matching follow-up
request and still checks that the root sees 50 tokens remaining after
the child uses its share.

This is test-only. Shared-budget behavior and the common response-mock
helper are unchanged.
2026-07-08 16:10:28 +01:00
charlesgong-openai
1ee0e9a949 Log plugin install failure subtypes (#31518)
## Summary

- classify plugin store and remote bundle I/O failures by privacy-safe
operation context
- include the normalized `sub_error_type` in structured plugin-install
failure warnings
- leave analytics events and schemas unchanged

## Why

Top-level error types such as `store_io` collapse multiple filesystem
operations, making logs harder to diagnose. The subtype identifies the
failed operation without logging paths or new exception details.

## Impact

Plugin-install warning logs gain an optional normalized subtype.
Analytics payloads, app-server APIs, and stored data are unchanged.

## Validation

- `just fmt`
- `just test -p codex-core-plugins -p codex-app-server` (1,235 passed;
14 failed and 6 timed out because local `sandbox-exec` was denied,
`test_stdio_server` was unavailable, or user-level skill fixtures were
loaded)
2026-07-08 11:06:45 -04:00
jif
f17a57b7d5 Stabilize remote compaction parity against dynamic skill catalogs (#31585)
## Why

The remote compaction parity test compares legacy and v2 sessions
created with separate temporary homes. Those sessions can discover
different model-visible skill catalogs, so the request comparison can
fail even when the compaction and service-tier behavior matches.

This is the most frequent retry-saved full-CI failure in the recent
JUnit history.

## What changed

Normalize only the contents of `<skills_instructions>` before comparing
the captured requests. The opening and closing tags remain in the
comparison, so the test still catches a missing or misplaced skills
block.

The service-tier, compacted input, follow-up request, and
replacement-history assertions are unchanged. A focused normalizer test
covers the new behavior.

## Scope

This is test-only. It does not change runtime compaction or skill
behavior. Exact skill-catalog rendering remains covered by the dedicated
skills tests.
2026-07-08 15:59:36 +01:00