October 2, 2026
AXIOM: Building a Practical Red-Teaming Platform for AI Systems
Part 1 โ The problem, the first architecture, and the verification principle that shaped everything after
By Nisarg Patel
7 min read
There is a gap in AI security tooling.
On one side are one-off scripts: a prompt injection test written for a particular model, a notebook used once during an assessment, or a small collection of payloads that never quite becomes reusable tooling.
On the other side are heavyweight research frameworks such as garak and PyRIT, designed to cover broad areas of AI security research.
Both approaches are useful. But neither quite matches the workflow of a security consultant who has been given authorization to test a specific AI product and needs to answer a practical question:
What can I safely test, what actually worked, and what can I hand back to the client?
That was the problem space AXIOM was built for.
AXIOM is an Adversarial AI Security Testing Platform for red-teaming LLM applications, broader ML models, and agentic systems with tool access. The simplest description we use is:
"Metasploit for AI."
Not because the two projects are technically identical, but because the idea is similar: give a security practitioner a unified target abstraction, reusable attack modules, and structured results instead of forcing every assessment to begin with a blank script.
AXIOM is not a SOC.
It is not an alert-correlation engine.
It does not contain an internal LLM that reasons over security events.
It is Python infrastructure whose job is to test AI systems under explicit authorization and produce structured findings about what happened.
Why A.X.I.O.M
The motivation was deliberately practical.
If a consultant is handed an AI system to assess, three things matter immediately:
- The target should be quick to configure. You should not need to write a new harness for every engagement.
- Authorization should be structural. It should not simply be a warning in a README telling the operator to behave responsibly.
- The output should resemble a security deliverable. A collection of raw responses is not particularly useful to a client.
Those requirements shaped the first architecture.
Not an elaborate pipeline.
Not an AI reasoning layer.
Just a small number of components with clearly defined responsibilities.
The First Architecture
The initial shape of AXIOM was deliberately simple:
scope.yaml
โ
authorization
โ
Target
โ
Recon
โ
attack module
โ
Finding
โ
JSON / Markdown reportscope.yaml
โ
authorization
โ
Target
โ
Recon
โ
attack module
โ
Finding
โ
JSON / Markdown reportEach part exists for a reason.
Scope comes first
scope.yaml defines what the platform is allowed to touch.
The configuration requires three legal acknowledgements:
authorized_testerresponsible_disclosureno_production_harm
All three must be enabled before modules can run.
Authorization is handled through ScopeManager.authorize(). A target that is not explicitly declared in scope does not simply generate a warning and continue. The platform refuses to proceed.
That became the first major design principle:
Scope is structural, not a warning.
This distinction matters in security tooling. The authorization boundary should exist in the execution path, not merely in the documentation.
One Target, many systems
The next problem was target diversity.
An LLM might be exposed through OpenAI or Anthropic. It might run locally through Ollama. A machine-learning assessment might involve a local model. An agentic system introduces another layer entirely.
The attack module should not have to care how the target is transported.
So AXIOM introduced a common Target abstraction.
Attack modules interact with the target through that interface rather than directly importing individual network or provider clients.
That led to the second principle:
One Target interface, many backends.
It also became one of the decisions that allowed the project to grow from LLM testing into ML and agentic testing without creating a separate architecture for every category.
Findings are data
An attack module should not return a paragraph and leave another component to figure out what it means.
AXIOM uses structured Finding objects containing information such as the module, category, attack type, severity, evidence, remediation, confidence, and reproducibility information.
The result can then be written to JSON for programmatic use or Markdown for a client-facing report.
That became the third principle:
Findings are structured, not strings.
The fourth principle came from failure
Security tooling has an uncomfortable failure mode: it is possible to produce impressive-looking reports full of findings that are not actually vulnerabilities.
A heuristic that is slightly too aggressive can turn noise into a security issue.
AXIOM therefore adopted another explicit principle:
Under-claim over over-claim.
When the evidence is insufficient, NO_FINDING is preferable to inventing confidence.
Unverified heuristic results are treated accordingly rather than being presented with the same certainty as verified findings.
That principle became particularly important once the tool was pointed at real models.
The hardest problem wasn't launching the attack
The interesting part of red-teaming AI systems is not necessarily generating an attack.
It is determining whether the attack actually worked.
A model can repeat words from a payload and make it look as though it disclosed something.
It can refuse an instruction while quoting the exact phrase the test was looking for.
It can produce a convincing-looking response that has no relationship to whether the underlying security condition occurred.
In other words:
Attack success is not self-reporting.
AXIOM gradually converged on a verification pattern:
The attack proposes the result. Evidence outside the attack determines whether the result is real.
The implementation differs across AI categories, but the principle remains the same.
For LLMs: Use a Canary
For system-prompt leakage testing, AXIOM can use a known canary marker placed in the target's configuration.
The marker should have no legitimate reason to appear in the model's response unless the protected content was actually disclosed.
That makes the result substantially stronger than simply searching for words that look like a system prompt.
A heuristic can suggest that something happened.
A canary can verify it.
For ML: Use Ground Truth
ML security testing introduced another version of the same idea.
AXIOM includes an owned NumPy MLP victim trained on the scikit-learn digits dataset, with vulnerable and hardened profiles.
The evaluation side knows which samples belong to the training set.
The attack modules do not.
That information is exposed through a separate GroundTruth mechanism so that an attack can produce its verdict independently, after which the harness can compare that verdict against the actual truth.
Again:
the attack does not get to define its own success.
For agents: Watch What Actually Happened
Agentic systems created another variation.
AXIOM's mock agent operates inside a sandbox rather than touching a real filesystem, mailbox, network, or shell.
The important evidence is therefore not what the agent says it would do.
It is what the tool log shows it actually attempted and executed.
If an attack claims that it caused an unauthorized tool call, the finding needs to correspond to an actual logged action.
The tool log becomes the oracle.
Three categories.
Three different verification mechanisms.
One underlying principle:
The attack proposes. The harness verifies.
Real targets exposed the weaknesses
This is where the architecture stopped being theoretical.
Running AXIOM against a real local model โ llama3 through Ollama โ exposed problems that were difficult to appreciate from code alone.
The first one was deceptively simple.
A refusal is not compliance
The prompt-injection module used a fixed marker such as CONFIRMED to determine whether the model accepted an override.
The original detection logic essentially asked:
Did the response contain
CONFIRMED?
Then llama3 produced a refusal along the lines of:
"I can't output CONFIRMED."
The marker was there.
The model had not complied.
AXIOM had nevertheless counted it as a successful attack.
That was a false positive.
The fix was to check for refusal patterns before crediting the marker. A shared looks_like_refusal helper was introduced so the same mistake would not be repeated across other LLM modules.
It was a small code change.
But it represented a much bigger lesson:
A string appearing in a response is not necessarily evidence of the condition you're testing.
An echo is not a leak
The second failure was even more revealing.
The system-prompt leakage test used a SYSTEM PROMPT: marker.
The initial detector saw the marker and assumed that whatever followed it was leaked content.
Then llama3 returned:
SYSTEM PROMPT: <user>SYSTEM PROMPT: <user>At first glance, that looked like a disclosure.
But <user> had come directly from the payload.
The model had simply echoed part of the attack input.
Nothing had been leaked.
The detector was measuring its own payload rather than the target's behavior.
The fix required the content after the marker to be substantive and not already present in the submitted payload. If it was merely an echo, it was discarded.
Again, the important lesson was not the particular conditional added to the detector.
It was that real target behavior exposed a flaw that static reasoning over the code would not have revealed.
That is one of the reasons testing the tester became such an important part of AXIOM's development.
From LLM Testing to ML to Agents
AXIOM did not start with all three categories.
It grew in phases.
Phase 0 โ Foundation
The first phase established the plumbing:
- scope enforcement
- target abstraction
- reconnaissance
- module contracts
- CLI
- JSON and Markdown reporting
- tests
There was no attack module yet.
The goal was to build the infrastructure that later attacks could reuse.
Phase 1 โ LLM Security Testing
The first attack modules were:
prompt_injectionjailbreak_evolutiondata_leakage
This was where the canary-based verification approach began taking shape.
jailbreak_evolution explored persona and framing manipulation through escalating generations of prompts, while data_leakage supported both verified canary testing and a more cautious heuristic fallback.
Phase 2 โ ML Security Testing
The next expansion added:
model_inversionmembership_inferenceadversarial_examplesdata_poisonin
Alongside them came the owned NumPy MLP victim.
This phase also made capability boundaries structural. A target without the capabilities required for a particular attack does not get a clever workaround. The module reports that it is blocked.
Phase 3 โ Agentic Security Testing
Then came:
tool_abusetask_hijack
and the sandboxed mock agent.
The same Target abstraction survived the transition.
There was no need to create an entirely separate transport architecture for agents.
That was the payoff from one of the earliest decisions in Phase 0.
What was still unfinished
At this point, AXIOM had a working core.
But it was not finished.
Phase 4 was planned around orchestration and unified scoring: chaining modules into strategies such as full, stealth, or category-specific scans, alongside a unified AI-CVSS-style severity model.
At the time represented by the source snapshots, that did not exist yet. Individual modules still determined severity using their own logic.
Phase 5 was planned as the service layer: a FastAPI backend, scan history, authentication, dashboard functionality, and PDF report export.
Those were future layers.
The working core was Phase 0 through Phase 3:
an authorized, modular, verification-first platform for red-teaming LLMs, ML models, and agentic systems.
And that distinction โ between what existed and what was planned โ became important too.
A security tool should not claim capabilities it does not have.
That is the same philosophy behind the fourth design principle.
Under-claim over over-claim.
The Foundation
Looking back at the first architecture, none of its individual components were particularly exotic.
A scope manager.
A target abstraction.
A reconnaissance engine.
Attack modules.
Structured findings.
Reports.
What mattered was how those pieces constrained one another.
The target could not simply be attacked outside scope.
The module did not own the transport layer.
The finding was not an unstructured paragraph.
A suspicious response was not automatically treated as proof.
And a planned feature was not described as though it already existed.
Those decisions created the foundation for everything that followed.
More importantly, the early failures established something that would continue through every later phase:
AI security testing is not only about generating better attacks. It is about building better evidence around the attacks.
That became one of the defining ideas behind AXIOM.
And once the platform moved beyond LLMs into ML and agentic systems, the same question kept returning in different forms:
How do you know the attack actually worked?
That question is where the next stage of AXIOM begins.