3 tools · listed in dataset order, no ranking
AI agents that test applications you are authorized to assess, verify vulnerabilities, and propose fixes.

Strix (OmniSecure, Inc.)
Strix: Open-Source AI Pentesting Agents for Authorized App Security Testing
Best for: Engineering and AppSec teams that want open-source, AI-driven pentesting of their own apps and APIs in the CLI or CI, with an option to move to a managed platform
Keygraph
Shannon: Open-Source White-Box AI Pentester for Web Apps and APIs
Best for: Engineering teams with access to both the source code and a staging copy of their web app or API who want open-source, evidence-backed security testing in CI

CodeMender: Google Cloud's AI Agent That Finds, Verifies, and Fixes Code Vulnerabilities
Best for: Security and platform teams already on Google Cloud who want an agent that verifies vulnerabilities and proposes tested patches, and who can join a preview program
Strix and Shannon test a running application (Shannon also reads its source code) and report issues they can validate. CodeMender works on source code: it verifies findings and drafts tested patches for developers to review. Many teams will want both kinds of coverage.
Active security testing tools can change application state. Their vendors say to run them only against systems you own or have written permission to test, ideally a staging environment with disposable data.
The open-source tools here are bring-your-own-model, so each scan incurs LLM costs, and deep runs can take hours. Managed platforms trade that for seat, per-test, or token-based pricing.
An AI agent that performs security work such as testing an application you are authorized to assess, validating whether a vulnerability is real, or drafting a fix, rather than only listing potential issues.
Yes. Strix is Apache-2.0 and Shannon is AGPL-3.0. Both run locally or in CI with your own model provider, and both vendors also offer paid hosted platforms.
No. Vendors say findings still need human review, and coverage is narrower than a full manual assessment. They are best used for continuous or pre-release checks alongside periodic human-led testing.