Mongoose turns autonomous agents loose on your models, RAG pipelines and tool-using agents. They fingerprint each target, plan the attack, adapt when they're refused and judge every response. Every finding maps to a control in your own framework, and confirmed ones become runtime blocking rules.


Tests the models, pipelines and agents you already run
AI-SPM is the practice of always knowing how exposed your AI systems are, and keeping that exposure under control. Cloud security posture management did this for infrastructure. AI-SPM does it for the models, RAG pipelines and agents now running in production, where the risks are prompt injection, data leakage, poisoned retrieval and agents with more access than they need.
Every model, RAG pipeline and agent you test sits in one view with its results side by side. Each agent carries an identity, an owner and a policy.
Agents attack each deployment, and batch runs across your fleet compare against a baseline, so a model or prompt change that reopens a hole shows up as a regression.
Every finding maps to the OWASP LLM Top 10 (2025) and to the controls in your own framework, so posture reads in the language your auditors use.
Confirmed findings become Vigilant rules. They watch in monitor mode first, then block once you've seen what they catch.
0
attack categories, OWASP-mapped
0
open-source frameworks, one runner
0
in-house Malagasy attack modules
0+
variations per Deep Scan
It can't react to how your model refuses, and it flags plenty of false positives. Coverage decays every time the model or system prompt changes.
misses: adaptive & multi-turn attacks
It blocks known shapes on the way in and out. Nobody measures its own bypass rate against your deployment.
misses: whether you're vulnerable at all
Useful posture data, but the risk rating has no adversarial evidence behind it.
misses: proof
Mongoose produces evidence. Every finding carries the full exchange, the judge's verdict and confidence, and the OWASP risk and control it implicates.
See howMongoose's agents share a blackboard, run under a hard budget, and persist every decision so a run can be audited or resumed. The reasoning agents use their own sandboxed credentials, never the model under test.
RAG, agent tool-use, multimodal, consumption and long-horizon attacks come from Malagasy, the in-house module I built to cover what the open-source tools don't reach.
Direct and indirect prompt injection, jailbreaks, latent injection.
Open-source frameworks
Leakage, training-data extraction, system-prompt disclosure.
Open-source frameworks
Bias, toxicity, misinformation, fabricated citations, harmful content.
Open-source frameworks
Base64, ROT13, atbash, ASCII art, obfuscation, prompt smuggling.
Open-source frameworks · Deep Scan
Knowledge poisoning, retrieval manipulation, context overflow, injection via retrieved docs.
Malagasy
Tool abuse, privilege escalation, agent hijacking, tool-chain exploitation.
Malagasy
Image injection, OCR bypass, steganography, cross-modal exploits.
Malagasy
Crescendo, refusal erosion, persona drift, context poisoning.
Malagasy · open-source frameworks
Token flooding, output amplification, recursive reasoning, wallet-drain simulation.
Malagasy (opt-in)
Open-source tools flag a lot of false positives. So Mongoose re-checks every hit they report before it reaches your results. A deterministic detector match stands; anything inferred goes to the judge, which confirms it, rejects it or sends it to review. Nine open-source frameworks and 24 Malagasy modules report into one schema.
Register your own control and threat catalog as a harness and sync it from GitHub. Every finding, from an agent run or a scanner, maps to the control it implicates, so an audit question gets answered from the same evidence as the red team.
via=scanner
The scanner names the control itself. The MCP compliance lens emits one per clause.
via=family-map
A scanner rule maps to its family in your harness, and from there to the controls, threats and OWASP entries it lists.
via=owasp-crosswalk
An adversarial finding's OWASP LLM category maps to every control whose crosswalk cites it. Coarser, and labelled that way.
Agent runs test live endpoints. Symphony reads source: MCP servers and agent skills, scanned deterministically, with every finding quoting file and line. A clean result means no statically decidable violation, not a safe repository.
Four lenses over an MCP server and its client configuration.
Skill and agent-pack definitions against the OWASP Agentic Skills Top 10.
When an agent calls a tool, you know which agent it was, which person launched it and whether it was allowed. Nothing blocks by default: rules start in monitor mode and only take effect once promoted, so onboarding a policy never breaks a run that worked yesterday.
verified and stripped by the Vigilant proxy before forwarding
One record per agent, whether it's a platform agent, MCP server, coding agent or custom. Each has an accountable owner and a tool manifest.
Every dispatched agent gets a signed, short-lived token naming the agent, the person who launched it, the run and its scope.
Plans, drills, tool calls and agent-to-agent messages are decided as auto, hold for a person, or block. New rules start in monitor mode.
Every entry carries the hash of the one before it. Filter the trail by agent or by run.
Standard discovery and JWKS endpoints, so your identity provider can trust Mongoose as an issuer. Keys rotate without breaking tokens in flight.
An identity attack category covers delegation abuse, confused deputy and owner impersonation (LLM06). Vigilant detects the same shapes in live traffic.
Not covered yet
Vigilant inspects production LLM traffic inline. A confirmed attack becomes a rule, the rule runs in monitor mode, and you promote it to enforcement once its false-positive rate holds up.
The built-in assistant reads your results with you. It explains a finding, separates confirmed from unproven, and tells you what to change.
Attach the evidence
Point it at a scan, a finding or a model, and it answers about that.
Grounded, or it says so
Claims about a result come from the attached context. When the context doesn't say, neither does it.
Fixes mapped to OWASP
Mitigations come tied to the OWASP LLM category, and a judge score is treated as a signal, not a verdict.
Advisory only
It answers questions. It doesn't launch scans or change rules.
In the app and the CLI
Ask Mongoose in the web app, Orion in the CLI. Bring your own Anthropic or OpenAI key.
Start with a scoped, authorized test against one real deployment. The findings make the case better than any slide.
One deployment you own, one written scope, one agentic run.
The same suite on every model and system-prompt change.
Vigilant in front of production traffic.
Run it in CI on every model version and system-prompt change, so "we tested it once" becomes "we know if the new version regressed." Targets OpenAI, Anthropic, Azure OpenAI and Microsoft Foundry, Google, Ollama and custom endpoints.
Tell us what you're running. A scoped test is the fastest way to see what the agents find on your own system.
A model, RAG pipeline or agent you own or have written permission to test.
Target, categories and budget agreed before any traffic is sent.
Fingerprint, adaptive attack plan and judged findings with the full exchange.
Confirmed, partial and refused findings, with likely false positives called out.