
Google · Google Mantis: Open-Source Agent Skills for Finding, Reproducing, and Patching Vulnerabilities
Google Mantis is an open-source (Apache-2.0) set of security-review skills, plus a reference harness built on Google's Agent Development Kit, that lets an AI coding agent threat-model a codebase, hunt for vulnerabilities, reproduce them in a sandbox, and write and re-attack patches. It suits security engineers and AppSec teams who want to run and adapt a vulnerability-research pipeline on their own infrastructure, with their own coding agent and models.
Best for
Security engineers and AppSec teams who can operate isolated sandboxes and want a free, adaptable pipeline that reproduces and patches vulnerabilities with their own coding agent and models
Not ideal for
Teams that want a supported, production-ready product, a hosted service, or pull-request reviews out of the box, and anyone without the capacity to review AI-generated findings or isolate code execution
Who it's for
Security engineers, AppSec teams, and security researchers who run their own coding agents and sandboxed infrastructure
Mantis is best read as a blueprint for AI vulnerability research rather than a finished scanner. Its design choices are the valuable part: grounding findings in sandboxed reproduction, having independent agents attack each patch, and calibrating severity against a rubric all address the false-positive and severity-inflation problems that make most AI code scanning hard to trust. The cost is that you own everything around it: sandbox isolation, model spend, run-to-run variance, expert triage, and any CI or ticketing integration. Teams with security engineers who want to build an internal pipeline will get the most from it; teams that would rather use a vendor-run product should look at CodeMender, which is in limited preview, or Codex Security.
Who should use it
Security engineering and AppSec teams that already run coding agents, can stand up isolated sandboxes, and want to build or tune their own vulnerability discovery and patching pipeline; also security researchers studying agentic bug-finding techniques.
Who should skip it
Teams that need a supported product or a hosted service, developers looking for drop-in pull-request security reviews, and anyone who cannot isolate AI-generated code execution or review findings with security expertise.
Open source
Free
Note: Mantis has no paid edition. Costs come from the model provider you configure and from sandbox infrastructure such as Compute Engine VMs. The harness supports wall-clock and token budgets per run, and resumed runs start with a fresh budget.
Available models
Building an internal AI bug-finding pipeline
The skills and published data contracts let a security team wrap each stage in its own deterministic harness, with sandboxing and reporting enforced by code rather than left to the model.
Triaging AI findings with evidence
Review, critic, and tiered reproduction stages separate findings that can be demonstrated in a sandbox from ones that cannot, so engineers start from evidence instead of model confidence.
Preventing repeat vulnerabilities
The `/mantis-advise` skill gives a coding agent the repository's threat model, past bug lineages, and verified patches while it writes new code, so fixed bug classes are less likely to return.
Google Mantis vs. Codex Security
Codex Security is OpenAI's application security product: it covers threat modeling, discovery, validation in isolation, and bounded patches, with a Codex plugin, an Apache-2.0 CLI and SDK, GitHub pull-request reviews, and a cloud service, but scans require Codex Security access. Mantis covers a similar loop as free skills and a reference harness that you run with your own coding agent and models, with no hosted service or PR reviewer. Codex Security fits teams that want a product inside OpenAI's tools, and Mantis fits teams that want to adapt the pipeline themselves.
Google Mantis vs. CodeMender
CodeMender is Google Cloud's managed find, verify, and fix agent, in limited public preview for select customers, billed per token at Gemini rates through a Google Cloud project. Mantis is a separate Google open-source project that you deploy and modify yourself, with your choice of models and sandboxes, and that Google labels not officially supported and intended for demonstration rather than production. CodeMender suits Google Cloud customers who can join its preview and want a managed path, and Mantis suits teams that want to own and customize the pipeline.
What is Google Mantis?
Mantis is an open-source toolkit from Google's security engineers that lets an AI coding agent run a vulnerability-research pipeline: it builds a threat model, hunts for bugs, filters false positives, reproduces findings in a sandbox, writes patches, re-attacks them, and scores the remaining risk. It ships as slash-command skills plus a reference harness built on Google's Agent Development Kit.
Is Mantis an official Google product?
It is published by Google in the google GitHub organization under the Apache-2.0 license, and Google Cloud's security blog presents it as the open-source core of the framework Google uses internally. The repository also says it is not an officially supported Google product, is intended for demonstration purposes, and is not meant for production use.
Is Mantis free?
Yes. The code is free under Apache-2.0. You pay for the models you run it with and for any sandbox infrastructure, such as Compute Engine VMs. The harness lets you cap wall-clock time and token use per run.
Which coding agents and models does Mantis work with?
The skills work with Gemini CLI, Antigravity CLI, Google ADK, and other coding agent frameworks. The reference harness supports Gemini models directly or through Vertex AI, Claude and GLM models through Vertex AI Model Garden, and any OpenAI-compatible endpoint, including vLLM, Ollama, or a LiteLLM proxy.
Is it safe to run Mantis on my machine?
Google warns that Mantis generates and executes code autonomously and should run only in isolated, restricted environments with no access to production systems, sensitive data, or internal networks. Its sandboxes run reproducers in networkless gVisor containers, microVMs, or hardened Compute Engine VMs, and Google requires a hardened VM for unattended runs.
Mantis vs CodeMender: what is the difference?
Both come from Google and both find, verify, and patch vulnerabilities. CodeMender is a managed Google Cloud agent in limited public preview with token-based Gemini pricing and a `cm` CLI. Mantis is a free, unsupported open-source toolkit that you run and adapt yourself with your own coding agent, models, and sandboxes.

OpenAI
Security and engineering teams that already use Codex or the OpenAI API and want an agent that validates findings and prepares reviewable fixes, from a one-off repository audit to per-PR checks in CI
Paid
Security and platform teams already on Google Cloud who want an agent that verifies vulnerabilities and proposes tested patches, and who can join a preview program
PaidMantis comes from Google's security engineering team, which describes it as the open-source core of the multi-agent framework Google Cloud uses internally to scan its own code; Google says a fuller version runs internally. It is not a scanner you point at a repository and leave. It ships as a set of slash-command skills that you install into a coding agent such as Gemini CLI, Antigravity CLI, or an ADK app (`npx skills add google/mantis`), plus a reference harness on the Agent Development Kit (`./run.sh path/to/code`) that chains the stages for you. A pass starts by mining version-control history for past security fixes, building a Markdown knowledge base and a threat model, and writing a targeted research plan. You can steer the plan in plain language with `--focus`, or have the harness design a custom agent graph for an objective. Research agents then sweep the code, and deduplication, review, and critic stages drop duplicates, false positives, and issues that cannot occur in a release build. Surviving findings are reproduced in tiers, from a unit micro-harness up to a full sandboxed service, inside networkless gVisor containers or microVMs, or on hardened ephemeral Compute Engine VMs in a private VPC without internet access. Confirmed findings can be chained into multi-step exploits. Patches are written by separate subagents and re-attacked by independent agents with boundary-mutated variants before a fix is marked verified. Every finding gets a 1-10 risk score against a rubric meant to stop the inflated severity ratings LLMs tend to give, and learnings feed the next pass and the `/mantis-advise` skill, which your coding agent can consult while writing new code. Google says a hierarchical summary tree cuts token overhead by more than 85% on large repositories. Models are your choice: Gemini directly or through Vertex AI, Claude and GLM through Vertex AI Model Garden, or any OpenAI-compatible endpoint such as vLLM or Ollama. The tradeoffs are significant. Google labels Mantis "not an officially supported Google product" and says it is for demonstration, not production. It generates and runs code autonomously, so it must run in isolated environments, and Google requires a hardened VM for unattended use. Results are non-deterministic and every finding needs review by a security expert. There is no hosted service, pull-request reviewer, or packaged CI integration, so teams build those around the harness themselves.
The reference harness gained a campaign planner with replanning and steering. The README shows steering the planner toward a bug class in plain language with `--focus`.
Google rewrote most of the reference harness, moving the multi-agent pipeline onto ADK workflows and hardening the boundary between the agents and the host.
Google Cloud's security blog published a getting-started guide for the open-source harness and introduced the mantis-advise skill for writing secure code with accumulated findings.
Google Cloud's CISO newsletter described Mantis as the multi-agent framework behind its internal AI code scanning and said its core skills are now open source.
Are you the founder? Claim this listing →