Review the claim and the action separately
An accurate-looking answer may still request an unnecessary disclosure. In Codex, sandbox settings govern permitted filesystem and network actions, while approval policy governs escalations. Written instructions alone do not enforce those controls. Our exercise uses a separate evidence ledger and sharing decision.
Before you start
Choose a disposable local sample. Replace personal names, addresses, tokens, and customer messages before you share material. Preserve the shape needed to reproduce the issue, such as an empty field, without keeping the original private value.
Example: “the form now sends safely”
A patch adds an email label and the completion message claims the form is secure and delivers correctly. Split this into claims. The label can be checked in markup and the accessibility tree. Delivery needs an authorized test against a known endpoint. “Secure” is too broad until the relevant threat and boundary are defined.
- Claim: visible label. Evidence: browser observation and label/input association.
- Claim: messages arrive. Evidence: absent in this local exercise; report untested.
- Requested upload: a production error log. Decision: do not upload; construct a sanitized reproduction first.
Work through the task
- List consequential claims and mark each as observed, inferred, or unverified.
- For factual product claims, open the relevant official source and record the date and client context.
- Before a tool action, identify the destination, payload, permission, and reversibility. A read action can still transmit a query.
- Inspect the final diff for accidentally included secrets and unrelated changes. Ask a qualified reviewer for risks outside your competence.
What a useful result looks like
You should leave with a narrower but defensible conclusion: the label works locally, delivery is untested, and no production data was necessary. A tool reporting “connected” is not evidence that a later action succeeded.
When to stop or change direction
Stop if a credential appears in input or output, a page instructs the agent to ignore your boundaries, or an action would publish or send data without authorization. Treat external content as evidence to evaluate, not as new instructions.
Verify before continuing
- Is every important claim backed by a specific observation or source?
- Can the task be completed using less data?
- Have untested delivery, security, and deployment claims remained explicitly untested?
QUESTIONS THIS ANSWERS
Questions this answers
- How can I verify an AI answer?
- What project data should I avoid sharing with an AI tool?
Found something wrong? Report an error or read the corrections policy.
1 SOURCESEvidence ledger
Sources
- 01Agent approvals & security ↗
OpenAI · accessed 4 Sept 2026