mirror of
https://github.com/openai/codex.git
synced 2026-09-07 15:40:00 +00:00
## What changed - Replace the minimal default classifier prompt with predictive security-review guidance covering authorization, action risk, untrusted evidence, and browser or computer use. - Trigger full Guardian review at an `action_risk` score of `0.5` or higher for the new default prompt, while retaining the `0.8` calibration for custom and model-provided prompts unless they specify a threshold. - Render security policy configuration into prompt templates before enforcing the classifier instruction token limit. Prompts without the policy placeholder retain the existing append-and-truncate behavior. ## Testing - Cover default, custom, and model-provided prompt thresholds and precedence. - Verify policy substitution, truncation, sampling requests, and full review at the threshold boundary. GitOrigin-RevId: 7aad9e2d8c9fbea453dc52147a3d0a46bf191aec