BetaEditorial Listing

Codex Security

OpenAI · Codex Security: OpenAI's AI Agent for Finding, Validating, and Fixing Code Vulnerabilities

Open Codex Security

Codex Security is OpenAI's application security agent. It builds a threat model of a repository, looks for vulnerabilities, tries to reproduce them in isolation, and prepares patches for you to review. It runs as a plugin in Codex, as an open-source CLI and TypeScript SDK, as a security reviewer on GitHub pull requests, and as Codex Security Cloud for connected GitHub repositories. It suits security and engineering teams whose ChatGPT workspace or account has Codex Security access and who want validated findings rather than a raw scanner list.

PricingPaid
Setupmedium
Runs onDesktop · Web · API
APIYes
Open sourceNo
DocsYes
CategorySecurity
Code SecurityVulnerability RemediationApplication SecurityThreat ModelingOpen Source CLICI/CDGitHubPreview

Best for

Security and engineering teams that already use Codex or the OpenAI API and want an agent that validates findings and prepares reviewable fixes, from a one-off repository audit to per-PR checks in CI

Not ideal for

Teams without Codex Security access or a qualifying ChatGPT plan, organizations that need deterministic, repeatable scanner output as compliance evidence, teams that want continuous cloud monitoring for code hosted outside GitHub, and anyone who needs a self-hosted service

Who it's for

Application security teams, security-minded engineering teams, and platform teams that run security checks in CI, especially organizations already on ChatGPT business plans or the OpenAI API

Capabilities

  • Standard scans that move through threat modeling, discovery, validation, attack-path analysis, and reporting, with each phase shown live in the Security workbench
  • Deep scans that run several workers and subagents for broader review, with limits on workers, discovery runs, and run time (up to 96 hours)
  • Validation that tries to reproduce likely vulnerabilities in an isolated environment and attaches the commands, logs, and results as evidence
  • Findings with severity, confidence, location, evidence, and remediation, plus a coverage report that marks each scan complete, partial, or unknown
  • Bounded fixes for approved findings, with a regression test where feasible, a verify-fix check that returns fixed, still_vulnerable, or inconclusive, and an advisory patch-risk assessment
  • Diff, working-tree, and path scans, a Git pre-commit hook that blocks high-severity findings, and CI examples for GitHub Actions and GitLab CI/CD with SARIF upload and severity thresholds
  • Bulk scans that discover repositories in a GitHub account or organization, or run a resumable campaign from a CSV inventory
  • Backlog triage of SARIF reports, GitHub code-scanning and Dependabot alerts, advisories, and Jira or Linear tickets against the current code
  • Scan history with cross-scan comparison (new, persisting, reopened, resolved), false-positive feedback that later scans take into account, and CSV, JSON, or SARIF export
  • Architecture documents, threat models, security policies, and custom instructions as scan context, plus an estimated USD cost limit per scan
  • Codex Security Review on GitHub pull requests, triggered when a PR opens, on every push, alongside Codex code review, or by commenting `@codex security review`
  • Codex Security Cloud (research preview): one-off repository scans and commit monitoring for connected GitHub repositories, with proposed patches opened as draft pull requests after review
  • TypeScript SDK for building scans, progress reporting, cost controls, and cancellation into your own tools

Limitations

  • The npm package is public, but running scans requires Codex Security access, and depending on the account and repository, full-repository scans may also require Trusted Access for Cyber verification
  • Codex Security Review is available on ChatGPT Pro, Business, Enterprise, and Edu, not Plus. Codex's plan comparison table lists Codex Security for connected GitHub repositories only in the Enterprise and Education column, so confirm what your workspace includes before relying on Cloud
  • No per-scan price is published. Pull-request reviews consume Codex allowance or ChatGPT credits, scans with an API key or a third-party provider are billed per token, and the `--max-cost` limit is an estimate, not a hard spending cap
  • Codex Security Cloud is a research preview, supports GitHub repositories only, and needs a compatible Codex cloud environment. OpenAI says scans of larger repositories may take several hours
  • Local and CI scans run with the permissions of the user or runner and do not pause for approval. Saved scan logs are not redacted and can contain source code or credentials
  • Findings can differ between runs of the same scan, and OpenAI says Codex Security complements SAST tools and manual security review rather than replacing them
  • The CLI needs Node.js 22.13+ (or 24 or 26), and scans, exports, and scan history also need Python 3.10+
  • Only the CLI, SDK, and local plugin workflows are open source. Codex Security Cloud and Codex Security Review run as OpenAI services, and no self-hosted cloud option is documented

Use cases

  • Running a first standard scan of a service before a release or an external security review
  • Adding a CI job that scans each pull request's changes, uploads SARIF to GitHub code scanning, and fails on high-severity findings
  • Scanning every non-archived repository in a GitHub organization as a resumable bulk campaign
  • Triaging a backlog of SAST, Dependabot, or bug-bounty findings to see which ones affect the code you actually run
  • Patching confirmed high and critical findings and opening a pull request with the verified fixes
  • Monitoring new commits on a critical GitHub repository with Codex Security Cloud

Our take

Codex Security's strongest idea is honesty about coverage. Every scan records what it reviewed and what it skipped, and OpenAI's own docs warn that a missing finding alone does not prove a fix worked, pointing users to reruns, direct rechecks, and a verify-fix check instead. That caution suits real security programs better than a confident list of findings. The range of surfaces is also useful, since the same scanner covers an ad-hoc desktop scan, a CI gate, a PR reviewer, and continuous cloud monitoring. The weak points are access and cost. Availability is gated and described inconsistently across OpenAI's pages, there is no per-scan price, and AI scans can vary between runs. Run it next to your existing scanners on a scoped service first, and check the coverage file before trusting a clean result.

Who should use it

Security teams and security-minded engineering teams that already work in Codex or with the OpenAI API, want validated findings with evidence, and plan to gate pull requests or audit many repositories without hiring more reviewers first.

Who should skip it

Teams without Codex Security access or an eligible ChatGPT plan, organizations that need deterministic scanner output or a self-hosted service, and teams whose code lives outside GitHub and who want managed continuous monitoring.

Strengths

  • Validates likely vulnerabilities and attaches evidence instead of returning a raw list of pattern matches
  • Coverage reports make it explicit what a scan did and did not review
  • Works from a desktop workbench, the terminal, CI, GitHub pull requests, and the cloud with the same scanner
  • The CLI, SDK, and local plugin workflows are Apache-2.0, and the CLI can use Amazon Bedrock, OpenRouter, or Fireworks models as well as OpenAI's
  • Fixes stay reviewable: nothing is applied or merged without a person choosing to

Weaknesses

  • Access is gated, and plan availability differs between OpenAI's own pages
  • No published per-scan price, and cost limits are estimates
  • Cloud is a research preview and GitHub-only
  • AI-driven results can vary between runs, so it does not replace deterministic scanners

Codex Security pricing

ChatGPT Pro, Business, Enterprise, and Edu

Included usage

  • Codex Security Review on GitHub pull requests (not available on Plus)
  • Reviews consume included Codex allowance or ChatGPT credits
  • Codex Security Cloud (research preview) where the workspace has access; Codex's plan table lists it under Enterprise and Education

CLI and SDK with an API key

Usage-based

  • Apache-2.0 package, free to install
  • Billed per token by OpenAI, Amazon Bedrock, OpenRouter, or Fireworks
  • Scans still require Codex Security access
  • Estimated per-scan cost limit with `--max-cost`

Note: OpenAI publishes no separate price or per-scan fee for Codex Security. Usage on ChatGPT plans draws on the Codex allowance and ChatGPT credits, and API-key scans are billed at the selected provider's token rates. The CLI reports token usage and estimated cost when available. Plan availability is described differently across OpenAI's pages: the Security Review docs list Pro, Business, Enterprise, and Edu, while Codex's plan comparison table lists Codex Security for connected GitHub repositories only under Enterprise and Education.

Technical specs

Available models

gpt-5.6-sol with xhigh reasoning effort (CLI default, and recommended for the plugin)Other OpenAI models your credentials can access, such as GPT-6.1 SolModels via Amazon Bedrock, OpenRouter, or Fireworks (CLI, explicit model required)

Where Codex Security excels

Gating pull requests on security findings

The CLI scans only the pull request's diff in GitHub Actions or GitLab CI/CD, uploads SARIF, and can fail the check above a severity you choose, while Codex Security Review can post security findings on GitHub pull requests automatically.

Auditing a whole GitHub organization

Bulk scans discover repositories or read a CSV inventory, run with configurable concurrency and retries, resume after interruption, and keep findings, coverage, and SARIF for each repository.

Working through an existing vulnerability backlog

Instead of running a new scan, Codex Security checks each existing finding from SARIF, Dependabot, advisories, or tickets against the current code and says whether the evidence supports action, shows the issue is not applicable, or needs more review.

Codex Security vs. competitors

Codex Security vs. CodeMender

CodeMender is Google Cloud's managed find, verify, and fix agent. It is in limited public preview, runs through a local `cm` CLI with Gemini models, and needs a Google Cloud project. Codex Security covers the same find, validate, and patch loop inside OpenAI's Codex products, with an open-source CLI and SDK, GitHub pull-request reviews, and a cloud service for GitHub repositories.

Codex Security vs. Shannon

Shannon is an AGPL-3.0 AI pentester that combines source analysis with attacks on a running staging app and reports only issues it can exploit. Codex Security works mainly from the code and repository context, needs no running deployment, and tries to reproduce issues in an isolated container when possible.

Codex Security vs. Strix

Strix is an Apache-2.0 AI pentesting tool whose agents test a local codebase, a repository, or a running app in a Docker sandbox and report validated findings, with a paid cloud platform for continuous testing, PR reviews, and autofix pull requests. Strix works with your own model keys and leans toward active testing of running targets. Codex Security starts from the source code with a threat model and fits most naturally for teams already using Codex and ChatGPT plans.

Frequently asked questions

What is Codex Security?

Codex Security is OpenAI's application security agent. It builds a threat model for a repository, searches for vulnerabilities, tries to reproduce likely issues in an isolated environment, and reports findings with severity, evidence, and remediation guidance. For findings you approve, it can prepare a focused patch and, where feasible, a regression test.

Is Codex Security the same as OpenAI Codex?

No. Codex is OpenAI's general coding agent. Codex Security is a separate security product with its own plugins, CLI, and SDK that run on Codex: the local plugin works in Codex in the ChatGPT desktop app and the Codex CLI, and Codex Security Cloud runs scans in Codex cloud. Codex's regular code review can also flag security issues, but Codex Security Review goes deeper on security-specific risks in pull requests.

How much does Codex Security cost?

OpenAI does not publish a separate price. Codex Security Review consumes the included Codex allowance or ChatGPT credits on eligible plans. The CLI can bill per token through an OpenAI API key or through Amazon Bedrock, OpenRouter, or Fireworks, and it reports token usage and estimated cost after a scan. You can set an estimated cost limit per scan with `--max-cost`, but it is not a hard cap.

Who can use Codex Security?

OpenAI's Security Review docs say Codex Security Review is available on ChatGPT Pro, Business, Enterprise, and Edu, not Plus. For Codex Security Cloud, Codex's plan comparison table currently lists Codex Security for connected GitHub repositories only under Enterprise and Education, although DevDay coverage reported Pro and Business access as well. The CLI package is public, but running scans requires Codex Security access, and some full-repository scans may need Trusted Access for Cyber verification. If access is unavailable, OpenAI's docs say to check with your workspace administrator.

Is Codex Security open source?

Partly. The `@openai/codex-security` CLI and TypeScript SDK, along with the local Codex Security plugin's workflows, are on GitHub at openai/codex-security under the Apache 2.0 license. Codex Security Cloud and Codex Security Review are OpenAI services, and running scans still requires Codex Security access.

Does Codex Security replace SAST tools?

No. OpenAI says Codex Security complements SAST: it adds reasoning about your specific code and automated validation, while deterministic scanners still provide broad, repeatable coverage. It can also triage findings those tools already produced, such as SARIF reports and Dependabot alerts.

Does Codex Security apply fixes automatically?

No. It proposes patches for findings you select, and in Codex Security Cloud you review the patch before creating a draft pull request. The CLI can commit verified patches and open a GitHub pull request, but only when you ask it to with `--patch` and `--create-pr`.

Codex Security vs CodeMender: what is the difference?

Both find, validate, and patch vulnerabilities in source code. CodeMender is Google Cloud's agent, in limited public preview for Google Cloud customers, with a local `cm` CLI and Gemini models. Codex Security runs on OpenAI's Codex products and ChatGPT plans, with an Apache-2.0 CLI and SDK, GitHub pull-request reviews, and a cloud service for GitHub repositories.

Integrations & fit

ChatGPT desktop appCodex CLICodex cloudGitHubGitHub ActionsGitHub code scanning (SARIF)DependabotGitLab CI/CDLinearJiraAmazon BedrockOpenRouterFireworks
Good fit forStartup / small team, Enterprise
Pricing modelPaid· Paid subscription required
See pricing on Codex Security →

Alternatives to consider

About Codex Security

Codex Security is OpenAI's dedicated application security product. It runs inside Codex as its own plugins, CLI, and pull-request reviewer, but instead of writing features it reviews code the way a security researcher would. A standard scan moves through threat modeling, discovery, validation, attack-path analysis, and reporting. A deep scan runs several workers for broader, longer review of a critical service or directory. Each finding records severity, confidence, location, evidence, and remediation guidance, and a separate coverage file lists what was reviewed, excluded, or deferred, so a scan with partial coverage is not mistaken for a clean result. Fixes are bounded: you choose which findings to patch, Codex Security prepares a focused change, and where it can, it adds a regression test that fails before the fix and passes after it. It does not apply patches on its own. There are four ways to use it. The Codex Security plugin adds a Security workbench (Scans, Findings, Repositories) to the ChatGPT desktop app and also works in the Codex CLI. The `@openai/codex-security` package, published with the plugin's workflows under Apache-2.0 in the openai/codex-security repository, provides a CLI and TypeScript SDK for local scans, bulk scans across a GitHub organization, a pre-commit hook, and CI jobs that upload SARIF and fail above a severity threshold. It can run with a ChatGPT sign-in, an OpenAI API key, or models from Amazon Bedrock, OpenRouter, or Fireworks. Codex Security Review adds security-focused reviews to GitHub pull requests. Codex Security Cloud, a separate plugin that OpenAI's docs describe as a research preview, scans connected GitHub repositories in ephemeral Codex cloud containers, either once or continuously as new commits land. Access and cost need the most attention. The npm package is public, but running scans requires Codex Security access, and some full-repository scans may also need Trusted Access for Cyber verification. Pull-request reviews consume Codex allowance or ChatGPT credits, API-key scans are billed per token, and OpenAI publishes no per-scan price. OpenAI positions it as a complement to deterministic SAST tools, not a replacement, and says results can vary between runs of the same scan.

Updates from Codex Security

New FeatureUpgraded Codex Security Cloud presented at OpenAI DevDay

At DevDay on September 29, 2026, OpenAI presented a major upgrade to Codex Security Cloud, which is now installed from the plugin marketplace on the web and in the desktop app. It scans connected GitHub repositories once or as new commits arrive and prepares fixes in the cloud. Launch coverage reported availability for Pro, Business, Enterprise, and Edu with Daybreak Blue model access included, while OpenAI's docs still describe Cloud as a research preview and Codex's plan table lists it under Enterprise and Education.

New FeaturePlugin 0.1.25 shows scan phases and keeps scan evidence

Codex Security plugin 0.1.25 shows standard scans advancing through threat modeling, discovery, validation, attack-path analysis, and reporting, stores retained scan evidence with saved results, and includes changed GitHub Actions workflow files and .cjs, .cts, .mts, and .tf files in reviews. Version 0.1.30 followed on September 24 with the same behavior.

New FeaturePlugin 0.1.22 adds fix verification

The verify-fix workflow checks whether an existing patch resolves a reported finding without changing repository files or issue trackers, and returns fixed, still_vulnerable, or inconclusive with supporting evidence. The remediation workflow also gained an investigation before patching and a review after it.

Are you the founder? Claim this listing →