
OpenAI · Codex Security: OpenAI's AI Agent for Finding, Validating, and Fixing Code Vulnerabilities
Codex Security is OpenAI's application security agent. It builds a threat model of a repository, looks for vulnerabilities, tries to reproduce them in isolation, and prepares patches for you to review. It runs as a plugin in Codex, as an open-source CLI and TypeScript SDK, as a security reviewer on GitHub pull requests, and as Codex Security Cloud for connected GitHub repositories. It suits security and engineering teams whose ChatGPT workspace or account has Codex Security access and who want validated findings rather than a raw scanner list.
Best for
Security and engineering teams that already use Codex or the OpenAI API and want an agent that validates findings and prepares reviewable fixes, from a one-off repository audit to per-PR checks in CI
Not ideal for
Teams without Codex Security access or a qualifying ChatGPT plan, organizations that need deterministic, repeatable scanner output as compliance evidence, teams that want continuous cloud monitoring for code hosted outside GitHub, and anyone who needs a self-hosted service
Who it's for
Application security teams, security-minded engineering teams, and platform teams that run security checks in CI, especially organizations already on ChatGPT business plans or the OpenAI API
Codex Security's strongest idea is honesty about coverage. Every scan records what it reviewed and what it skipped, and OpenAI's own docs warn that a missing finding alone does not prove a fix worked, pointing users to reruns, direct rechecks, and a verify-fix check instead. That caution suits real security programs better than a confident list of findings. The range of surfaces is also useful, since the same scanner covers an ad-hoc desktop scan, a CI gate, a PR reviewer, and continuous cloud monitoring. The weak points are access and cost. Availability is gated and described inconsistently across OpenAI's pages, there is no per-scan price, and AI scans can vary between runs. Run it next to your existing scanners on a scoped service first, and check the coverage file before trusting a clean result.
Who should use it
Security teams and security-minded engineering teams that already work in Codex or with the OpenAI API, want validated findings with evidence, and plan to gate pull requests or audit many repositories without hiring more reviewers first.
Who should skip it
Teams without Codex Security access or an eligible ChatGPT plan, organizations that need deterministic scanner output or a self-hosted service, and teams whose code lives outside GitHub and who want managed continuous monitoring.
ChatGPT Pro, Business, Enterprise, and Edu
Included usage
CLI and SDK with an API key
Usage-based
Note: OpenAI publishes no separate price or per-scan fee for Codex Security. Usage on ChatGPT plans draws on the Codex allowance and ChatGPT credits, and API-key scans are billed at the selected provider's token rates. The CLI reports token usage and estimated cost when available. Plan availability is described differently across OpenAI's pages: the Security Review docs list Pro, Business, Enterprise, and Edu, while Codex's plan comparison table lists Codex Security for connected GitHub repositories only under Enterprise and Education.
Available models
Gating pull requests on security findings
The CLI scans only the pull request's diff in GitHub Actions or GitLab CI/CD, uploads SARIF, and can fail the check above a severity you choose, while Codex Security Review can post security findings on GitHub pull requests automatically.
Auditing a whole GitHub organization
Bulk scans discover repositories or read a CSV inventory, run with configurable concurrency and retries, resume after interruption, and keep findings, coverage, and SARIF for each repository.
Working through an existing vulnerability backlog
Instead of running a new scan, Codex Security checks each existing finding from SARIF, Dependabot, advisories, or tickets against the current code and says whether the evidence supports action, shows the issue is not applicable, or needs more review.
Codex Security vs. CodeMender
CodeMender is Google Cloud's managed find, verify, and fix agent. It is in limited public preview, runs through a local `cm` CLI with Gemini models, and needs a Google Cloud project. Codex Security covers the same find, validate, and patch loop inside OpenAI's Codex products, with an open-source CLI and SDK, GitHub pull-request reviews, and a cloud service for GitHub repositories.
Codex Security vs. Shannon
Shannon is an AGPL-3.0 AI pentester that combines source analysis with attacks on a running staging app and reports only issues it can exploit. Codex Security works mainly from the code and repository context, needs no running deployment, and tries to reproduce issues in an isolated container when possible.
Codex Security vs. Strix
Strix is an Apache-2.0 AI pentesting tool whose agents test a local codebase, a repository, or a running app in a Docker sandbox and report validated findings, with a paid cloud platform for continuous testing, PR reviews, and autofix pull requests. Strix works with your own model keys and leans toward active testing of running targets. Codex Security starts from the source code with a threat model and fits most naturally for teams already using Codex and ChatGPT plans.
What is Codex Security?
Codex Security is OpenAI's application security agent. It builds a threat model for a repository, searches for vulnerabilities, tries to reproduce likely issues in an isolated environment, and reports findings with severity, evidence, and remediation guidance. For findings you approve, it can prepare a focused patch and, where feasible, a regression test.
Is Codex Security the same as OpenAI Codex?
No. Codex is OpenAI's general coding agent. Codex Security is a separate security product with its own plugins, CLI, and SDK that run on Codex: the local plugin works in Codex in the ChatGPT desktop app and the Codex CLI, and Codex Security Cloud runs scans in Codex cloud. Codex's regular code review can also flag security issues, but Codex Security Review goes deeper on security-specific risks in pull requests.
How much does Codex Security cost?
OpenAI does not publish a separate price. Codex Security Review consumes the included Codex allowance or ChatGPT credits on eligible plans. The CLI can bill per token through an OpenAI API key or through Amazon Bedrock, OpenRouter, or Fireworks, and it reports token usage and estimated cost after a scan. You can set an estimated cost limit per scan with `--max-cost`, but it is not a hard cap.
Who can use Codex Security?
OpenAI's Security Review docs say Codex Security Review is available on ChatGPT Pro, Business, Enterprise, and Edu, not Plus. For Codex Security Cloud, Codex's plan comparison table currently lists Codex Security for connected GitHub repositories only under Enterprise and Education, although DevDay coverage reported Pro and Business access as well. The CLI package is public, but running scans requires Codex Security access, and some full-repository scans may need Trusted Access for Cyber verification. If access is unavailable, OpenAI's docs say to check with your workspace administrator.
Is Codex Security open source?
Partly. The `@openai/codex-security` CLI and TypeScript SDK, along with the local Codex Security plugin's workflows, are on GitHub at openai/codex-security under the Apache 2.0 license. Codex Security Cloud and Codex Security Review are OpenAI services, and running scans still requires Codex Security access.
Does Codex Security replace SAST tools?
No. OpenAI says Codex Security complements SAST: it adds reasoning about your specific code and automated validation, while deterministic scanners still provide broad, repeatable coverage. It can also triage findings those tools already produced, such as SARIF reports and Dependabot alerts.
Does Codex Security apply fixes automatically?
No. It proposes patches for findings you select, and in Codex Security Cloud you review the patch before creating a draft pull request. The CLI can commit verified patches and open a GitHub pull request, but only when you ask it to with `--patch` and `--create-pr`.
Codex Security vs CodeMender: what is the difference?
Both find, validate, and patch vulnerabilities in source code. CodeMender is Google Cloud's agent, in limited public preview for Google Cloud customers, with a local `cm` CLI and Gemini models. Codex Security runs on OpenAI's Codex products and ChatGPT plans, with an Apache-2.0 CLI and SDK, GitHub pull-request reviews, and a cloud service for GitHub repositories.

Security and platform teams already on Google Cloud who want an agent that verifies vulnerabilities and proposes tested patches, and who can join a preview program
PaidKeygraph
Engineering teams with access to both the source code and a staging copy of their web app or API who want open-source, evidence-backed security testing in CI
Freemium
Strix (OmniSecure, Inc.)
Engineering and AppSec teams that want open-source, AI-driven pentesting of their own apps and APIs in the CLI or CI, with an option to move to a managed platform
FreemiumCodex Security is OpenAI's dedicated application security product. It runs inside Codex as its own plugins, CLI, and pull-request reviewer, but instead of writing features it reviews code the way a security researcher would. A standard scan moves through threat modeling, discovery, validation, attack-path analysis, and reporting. A deep scan runs several workers for broader, longer review of a critical service or directory. Each finding records severity, confidence, location, evidence, and remediation guidance, and a separate coverage file lists what was reviewed, excluded, or deferred, so a scan with partial coverage is not mistaken for a clean result. Fixes are bounded: you choose which findings to patch, Codex Security prepares a focused change, and where it can, it adds a regression test that fails before the fix and passes after it. It does not apply patches on its own. There are four ways to use it. The Codex Security plugin adds a Security workbench (Scans, Findings, Repositories) to the ChatGPT desktop app and also works in the Codex CLI. The `@openai/codex-security` package, published with the plugin's workflows under Apache-2.0 in the openai/codex-security repository, provides a CLI and TypeScript SDK for local scans, bulk scans across a GitHub organization, a pre-commit hook, and CI jobs that upload SARIF and fail above a severity threshold. It can run with a ChatGPT sign-in, an OpenAI API key, or models from Amazon Bedrock, OpenRouter, or Fireworks. Codex Security Review adds security-focused reviews to GitHub pull requests. Codex Security Cloud, a separate plugin that OpenAI's docs describe as a research preview, scans connected GitHub repositories in ephemeral Codex cloud containers, either once or continuously as new commits land. Access and cost need the most attention. The npm package is public, but running scans requires Codex Security access, and some full-repository scans may also need Trusted Access for Cyber verification. Pull-request reviews consume Codex allowance or ChatGPT credits, API-key scans are billed per token, and OpenAI publishes no per-scan price. OpenAI positions it as a complement to deterministic SAST tools, not a replacement, and says results can vary between runs of the same scan.
At DevDay on September 29, 2026, OpenAI presented a major upgrade to Codex Security Cloud, which is now installed from the plugin marketplace on the web and in the desktop app. It scans connected GitHub repositories once or as new commits arrive and prepares fixes in the cloud. Launch coverage reported availability for Pro, Business, Enterprise, and Edu with Daybreak Blue model access included, while OpenAI's docs still describe Cloud as a research preview and Codex's plan table lists it under Enterprise and Education.
Codex Security plugin 0.1.25 shows standard scans advancing through threat modeling, discovery, validation, attack-path analysis, and reporting, stores retained scan evidence with saved results, and includes changed GitHub Actions workflow files and .cjs, .cts, .mts, and .tf files in reviews. Version 0.1.30 followed on September 24 with the same behavior.
The verify-fix workflow checks whether an existing patch resolves a reported finding without changing repository files or issue trackers, and returns fixed, still_vulnerable, or inconclusive with supporting evidence. The remediation workflow also gained an investigation before patching and a review after it.
Are you the founder? Claim this listing →