Commit Graph

4372 Commits

Author SHA1 Message Date
Ahmed Ibrahim
75499a381b codex: fix Bazel insta snapshot paths (#13593) 2026-03-07 10:42:25 -08:00
Ahmed Ibrahim
bddc6cd53e codex: align guardian snapshot metadata 2026-03-07 10:26:19 -08:00
Ahmed Ibrahim
7456d28a7a codex: include core snapshots in bazel tests 2026-03-07 10:25:14 -08:00
Ahmed Ibrahim
406b5b386c codex: fix non-linux guardian popup snapshot 2026-03-07 10:21:29 -08:00
Ahmed Ibrahim
da2ddb4bf0 codex: stabilize guardian snapshots on linux 2026-03-07 10:19:58 -08:00
Ahmed Ibrahim
2260beb57f codex: restore guardian popup snapshot 2026-03-07 10:18:57 -08:00
Ahmed Ibrahim
dba2ee22dc codex: fix guardian popup snapshot 2026-03-07 10:13:34 -08:00
Ahmed Ibrahim
ea0e790813 codex: stabilize guardian permissions validation test 2026-03-07 09:54:59 -08:00
Ahmed Ibrahim
c42121a9f1 codex: fix guardian CI drift on PR #13593 2026-03-07 09:39:19 -08:00
Ahmed Ibrahim
f0cb95be20 Merge remote-tracking branch 'origin/main' into dev/pr13593-lint-fix 2026-03-07 09:10:33 -08:00
jif-oai
b9a2e40001 tmp: drop artifact skills (#13851) 2026-03-07 18:04:05 +01:00
Ahmed Ibrahim
e6f39a2d94 codex: fix rmcp lint regression 2026-03-07 08:52:55 -08:00
Ahmed Ibrahim
78fb500503 Merge branch 'main' into codex/flaky-test-stabilization-3 2026-03-07 08:47:37 -08:00
Ahmed Ibrahim
9c5049c21a codex: clean flaky test stabilization PR 2026-03-07 08:36:14 -08:00
Charley Cunningham
e84ee33cc0 Add guardian approval MVP (#13692)
## Summary
- add the guardian reviewer flow for `on-request` approvals in command,
patch, sandbox-retry, and managed-network approval paths
- keep guardian behind `features.guardian_approval` instead of exposing
a public `approval_policy = guardian` mode
- route ordinary `OnRequest` approvals to the guardian subagent when the
feature is enabled, without changing the public approval-mode surface

## Public model
- public approval modes stay unchanged
- guardian is enabled via `features.guardian_approval`
- when that feature is on, `approval_policy = on-request` keeps the same
approval boundaries but sends those approval requests to the guardian
reviewer instead of the user
- `/experimental` only persists the feature flag; it does not rewrite
`approval_policy`
- CLI and app-server no longer expose a separate `guardian` approval
mode in this PR

## Guardian reviewer
- the reviewer runs as a normal subagent and reuses the existing
subagent/thread machinery
- it is locked to a read-only sandbox and `approval_policy = never`
- it does not inherit user/project exec-policy rules
- it prefers `gpt-5.4` when the current provider exposes it, otherwise
falls back to the parent turn's active model
- it fail-closes on timeout, startup failure, malformed output, or any
other review error
- it currently auto-approves only when `risk_score < 80`

## Review context and policy
- guardian mirrors `OnRequest` approval semantics rather than
introducing a separate approval policy
- explicit `require_escalated` requests follow the same approval surface
as `OnRequest`; the difference is only who reviews them
- managed-network allowlist misses that enter the approval flow are also
reviewed by guardian
- the review prompt includes bounded recent transcript history plus
recent tool call/result evidence
- transcript entries and planned-action strings are truncated with
explicit `<guardian_truncated ... />` markers so large payloads stay
bounded
- apply-patch reviews include the full patch content (without
duplicating the structured `changes` payload)
- the guardian request layout is snapshot-tested using the same
model-visible Responses request formatter used elsewhere in core

## Guardian network behavior
- the guardian subagent inherits the parent session's managed-network
allowlist when one exists, so it can use the same approved network
surface while reviewing
- exact session-scoped network approvals are copied into the guardian
session with protocol/port scope preserved
- those copied approvals are now seeded before the guardian's first turn
is submitted, so inherited approvals are available during any immediate
review-time checks

## Out of scope / follow-ups
- the sandbox-permission validation split was pulled into a separate PR
and is not part of this diff
- a future follow-up can enable `serde_json` preserve-order in
`codex-core` and then simplify the guardian action rendering further

---------

Co-authored-by: Codex <noreply@openai.com>
2026-03-07 05:40:10 -08:00
Ahmed Ibrahim
faf223c4ae codex: validate CI stability (#13593) 2026-03-07 04:40:19 -08:00
Ahmed Ibrahim
762d419a48 codex: validate CI stability (#13593) 2026-03-07 03:56:55 -08:00
jif-oai
cf143bf71e feat: simplify DB further (#13771) 2026-03-07 03:48:36 -08:00
Ahmed Ibrahim
38354d89f0 codex: validate CI stability (#13593) 2026-03-07 03:27:33 -08:00
Ahmed Ibrahim
e064e7035b codex: validate CI stability (#13593) 2026-03-07 03:05:51 -08:00
Ahmed Ibrahim
c0b50b4ddc codex: add missing initialize forwarding hook (#13593) 2026-03-07 02:43:59 -08:00
Ahmed Ibrahim
0de39cd7ce codex: order websocket initialize readiness after handshake (#13593) 2026-03-07 02:39:20 -08:00
Ahmed Ibrahim
93af8e0d57 codex: validate flaky stabilization streak (#13593) 2026-03-07 02:10:32 -08:00
Ahmed Ibrahim
3b79f58849 codex: stabilize shell serialization duration test (#13593) 2026-03-07 01:46:25 -08:00
Ahmed Ibrahim
488a602833 codex: stabilize abort history test (#13593) 2026-03-07 01:33:48 -08:00
Ahmed Ibrahim
72cf2819be codex: normalize schema fixture TS paths (#13593) 2026-03-07 01:18:23 -08:00
Ahmed Ibrahim
e64133e403 codex: fix schema fixture compile regression (#13593) 2026-03-07 00:55:47 -08:00
Ahmed Ibrahim
c00078dd91 codex: reduce flaky test timeout pressure (#13593) 2026-03-07 00:50:42 -08:00
Ahmed Ibrahim
e951d6167f codex: validate flaky CI streak (5/5) (#13593) 2026-03-07 00:08:58 -08:00
Michael Bolin
5ceff6588e safety: honor filesystem policy carveouts in apply_patch (#13445)
## Why

`apply_patch` safety approval was still checking writable paths through
the legacy `SandboxPolicy` projection.

That can hide explicit `none` carveouts when a split filesystem policy
projects back to compatibility `ExternalSandbox`, which leaves one more
approval path that can auto-approve writes inside paths that are
intentionally blocked.

## What changed

- passed `turn.file_system_sandbox_policy` into `assess_patch_safety`
- changed writable-path checks to derive effective access from
`FileSystemSandboxPolicy` instead of the legacy `SandboxPolicy`
- made those checks reject explicit unreadable roots before considering
broad write access or writable roots
- added regression coverage showing that an `ExternalSandbox`
compatibility projection still asks for approval when the split
filesystem policy blocks a subpath

## Verification

- `cargo test -p codex-core safety::tests::`
- `cargo test -p codex-core test_sandbox_config_parsing`
- `cargo clippy -p codex-core --all-targets -- -D warnings`

---
[//]: # (BEGIN SAPLING FOOTER)
Stack created with [Sapling](https://sapling-scm.com). Best reviewed
with [ReviewStack](https://reviewstack.dev/openai/codex/pull/13445).
* #13453
* #13452
* #13451
* #13449
* #13448
* __->__ #13445
* #13440
* #13439

---------

Co-authored-by: viyatb-oai <viyatb@openai.com>
2026-03-07 08:01:08 +00:00
Ahmed Ibrahim
2b82a61f7e codex: validate flaky CI streak (4/5) (#13593) 2026-03-06 23:48:01 -08:00
Eric Traut
8df4d9b3b2 Add Fast mode status-line indicator (#13670)
Addresses feature request #13660

Adds new option to `/statusline` so the status line can display "fast
on" or "fast off"

Summary
- introduce a `FastMode` status-line item so `/statusline` can render
explicit `Fast on`/`Fast off` text for the service tier
- wire the item into the picker metadata and resolve its string from
`ChatWidget` without adding any unrelated `thread-name` logic or storage
changes
- ensure the refresh paths keep the cached footer in sync when the
service tier (fast mode) changes

Testing
- Manually tested

Here's what it looks like when enabled:

<img width="366" height="75" alt="image"
src="https://github.com/user-attachments/assets/7f992d2b-6dab-49ed-aa43-ad496f56f193"
/>
2026-03-07 00:42:08 -07:00
Ahmed Ibrahim
a5e13e321b codex: validate flaky CI streak (3/5) (#13593) 2026-03-06 23:21:57 -08:00
iceweasel-oai
4b4f61d379 app-server: require absolute cwd for windowsSandbox/setupStart (#13833)
## Summary
- require windowsSandbox/setupStart.cwd to be an AbsolutePathBuf
- reject relative cwd values at request parsing instead of normalizing
them later in the setup flow
- add RPC-layer coverage for relative cwd rejection and update the
checked-in protocol schemas/docs

## Why
windowsSandbox/setupStart was carrying the client-provided cwd as a raw
PathBuf for command_cwd while config derivation normalized the same
value into an absolute policy_cwd.

That left room for relative-path ambiguity in the setup path, especially
for inputs like cwd: "repo". Making the RPC accept only absolute paths
removes that split entirely: the handler now receives one
already-validated absolute path and uses it for both config derivation
and setup.

This keeps the trust model unchanged. Trusted clients could already
choose the session cwd; this change is only about making the setup RPC
reject relative paths so command_cwd and policy_cwd cannot diverge.

## Testing
- cargo test -p codex-app-server windows_sandbox_setup (run locally by
user)
- cargo test -p codex-app-server-protocol windows_sandbox (run locally
by user)
2026-03-06 22:47:08 -08:00
Celia Chen
b0ce16c47a fix(core): respect reject policy by approval source for skill scripts (#13816)
## Summary
- distinguish reject-policy handling for prefix-rule approvals versus
sandbox approvals in Unix shell escalation
- keep prompting for skill-script execution when `rules=true` but
`sandbox_approval=false`, instead of denying the command up front
- add regression coverage for both skill-script reject-policy paths in
`codex-rs/core/tests/suite/skill_approval.rs`
2026-03-06 21:43:14 -08:00
Ahmed Ibrahim
b5208d7979 codex: validate flaky CI streak (2/5) (#13593) 2026-03-06 21:38:43 -08:00
Ahmed Ibrahim
518c9a7ccf codex: fix rmcp pid-file race (#13593) 2026-03-06 21:15:48 -08:00
Ahmed Ibrahim
12c68ddc19 codex: shrink flaky protocol export test (#13593) 2026-03-06 21:03:34 -08:00
Ahmed Ibrahim
a9406ce8e8 codex: satisfy realtime startup context clippy (#13593) 2026-03-06 20:34:20 -08:00
Ahmed Ibrahim
484668e073 codex: fix realtime startup context close race (#13593) 2026-03-06 20:28:25 -08:00
Ahmed Ibrahim
ce981d77d2 codex: fix flaky realtime startup context test (#13593) 2026-03-06 20:14:50 -08:00
Ahmed Ibrahim
427880981c codex: validate flaky test stabilization (#13593) [4/5] 2026-03-06 20:03:17 -08:00
Michael Bolin
b52c18e414 protocol: derive effective file access from filesystem policies (#13440)
## Why

`#13434` and `#13439` introduce split filesystem and network policies,
but the only code that could answer basic filesystem questions like "is
access effectively unrestricted?" or "which roots are readable and
writable for this cwd?" still lived on the legacy `SandboxPolicy` path.

That would force later backends to either keep projecting through
`SandboxPolicy` or duplicate path-resolution logic. This PR moves those
queries onto `FileSystemSandboxPolicy` itself so later runtime and
platform changes can consume the split policy directly.

## What changed

- added `FileSystemSandboxPolicy` helpers for full-read/full-write
checks, platform-default reads, readable roots, writable roots, and
explicit unreadable roots resolved against a cwd
- added a shared helper for the default read-only carveouts under
writable roots so the legacy and split-policy paths stay aligned
- added protocol coverage for full-access detection and derived
readable, writable, and unreadable roots

## Verification

- added protocol coverage in `protocol/src/protocol.rs` and
`protocol/src/permissions.rs` for full-root access and derived
filesystem roots
- verified the current PR state with `just clippy`




---
[//]: # (BEGIN SAPLING FOOTER)
Stack created with [Sapling](https://sapling-scm.com). Best reviewed
with [ReviewStack](https://reviewstack.dev/openai/codex/pull/13440).
* #13453
* #13452
* #13451
* #13449
* #13448
* #13445
* __->__ #13440
* #13439

---------

Co-authored-by: viyatb-oai <viyatb@openai.com>
2026-03-07 03:49:29 +00:00
Ahmed Ibrahim
38741c5c78 codex: validate flaky test stabilization (#13593) [3/5] 2026-03-06 19:39:21 -08:00
Ahmed Ibrahim
82cc8397d4 codex: validate flaky test stabilization (#13593) [2/5] 2026-03-06 19:16:29 -08:00
Ahmed Ibrahim
3f393b65fb codex: fix flaky shell serialization timeout (#13593) 2026-03-06 18:45:46 -08:00
Ahmed Ibrahim
56bf69c219 codex: validate flaky test stabilization (#13593) [2/5] 2026-03-06 18:33:10 -08:00
Michael Bolin
22ac6b9aaa sandboxing: plumb split sandbox policies through runtime (#13439)
## Why

`#13434` introduces split `FileSystemSandboxPolicy` and
`NetworkSandboxPolicy`, but the runtime still made most execution-time
sandbox decisions from the legacy `SandboxPolicy` projection.

That projection loses information about combinations like unrestricted
filesystem access with restricted network access. In practice, that
means the runtime can choose the wrong platform sandbox behavior or set
the wrong network-restriction environment for a command even when config
has already separated those concerns.

This PR carries the split policies through the runtime so sandbox
selection, process spawning, and exec handling can consult the policy
that actually matters.

## What changed

- threaded `FileSystemSandboxPolicy` and `NetworkSandboxPolicy` through
`TurnContext`, `ExecRequest`, sandbox attempts, shell escalation state,
unified exec, and app-server exec overrides
- updated sandbox selection in `core/src/sandboxing/mod.rs` and
`core/src/exec.rs` to key off `FileSystemSandboxPolicy.kind` plus
`NetworkSandboxPolicy`, rather than inferring behavior only from the
legacy `SandboxPolicy`
- updated process spawning in `core/src/spawn.rs` and the platform
wrappers to use `NetworkSandboxPolicy` when deciding whether to set
`CODEX_SANDBOX_NETWORK_DISABLED`
- kept additional-permissions handling and legacy `ExternalSandbox`
compatibility projections aligned with the split policies, including
explicit user-shell execution and Windows restricted-token routing
- updated callers across `core`, `app-server`, and `linux-sandbox` to
pass the split policies explicitly

## Verification

- added regression coverage in `core/tests/suite/user_shell_cmd.rs` to
verify `RunUserShellCommand` does not inherit
`CODEX_SANDBOX_NETWORK_DISABLED` from the active turn
- added coverage in `core/src/exec.rs` for Windows restricted-token
sandbox selection when the legacy projection is `ExternalSandbox`
- updated Linux sandbox coverage in
`linux-sandbox/tests/suite/landlock.rs` to exercise the split-policy
exec path
- verified the current PR state with `just clippy`




---
[//]: # (BEGIN SAPLING FOOTER)
Stack created with [Sapling](https://sapling-scm.com). Best reviewed
with [ReviewStack](https://reviewstack.dev/openai/codex/pull/13439).
* #13453
* #13452
* #13451
* #13449
* #13448
* #13445
* #13440
* __->__ #13439

---------

Co-authored-by: viyatb-oai <viyatb@openai.com>
2026-03-07 02:30:21 +00:00
Ahmed Ibrahim
c38f0516c6 codex: fix flaky schema fixture timeout (#13593) 2026-03-06 18:08:01 -08:00
viyatb-oai
25fa974166 fix: support managed network allowlist controls (#12752)
## Summary
- treat `requirements.toml` `allowed_domains` and `denied_domains` as
managed network baselines for the proxy
- in restricted modes by default, build the effective runtime policy
from the managed baseline plus user-configured allowlist and denylist
entries, so common hosts can be pre-approved without blocking later user
expansion
- add `experimental_network.managed_allowed_domains_only = true` to pin
the effective allowlist to managed entries, ignore user allowlist
additions, and hard-deny non-managed domains without prompting
- apply `managed_allowed_domains_only` anywhere managed network
enforcement is active, including full access, while continuing to
respect denied domains from all sources
- add regression coverage for merged-baseline behavior, managed-only
behavior, and full-access managed-only enforcement

## Behavior
Assuming `requirements.toml` defines both
`experimental_network.allowed_domains` and
`experimental_network.denied_domains`.

### Default mode
- By default, the effective allowlist is
`experimental_network.allowed_domains` plus user or persisted allowlist
additions.
- By default, the effective denylist is
`experimental_network.denied_domains` plus user or persisted denylist
additions.
- Allowlist misses can go through the network approval flow.
- Explicit denylist hits and local or private-network blocks are still
hard-denied.
- When `experimental_network.managed_allowed_domains_only = true`, only
managed `allowed_domains` are respected, user allowlist additions are
ignored, and non-managed domains are hard-denied without prompting.
- Denied domains continue to be respected from all sources.

### Full access
- With managed requirements present, the effective allowlist is pinned
to `experimental_network.allowed_domains`.
- With managed requirements present, the effective denylist is pinned to
`experimental_network.denied_domains`.
- There is no allowlist-miss approval path in full access.
- Explicit denylist hits are hard-denied.
- `experimental_network.managed_allowed_domains_only = true` now also
applies in full access, so managed-only behavior remains in effect
anywhere managed network enforcement is active.
2026-03-06 17:52:54 -08:00