October 9, 2026
Red-Teaming Your Own MCP Server: A Checklist and a Test Harness
43% of tested servers had command injection. Go find yours first.

By Iamarbaz
4 min read
This is the companion to my piece on MCP tool poisoning. That one was about what a malicious server does to you. This one is the other side: you have built and shipped an MCP server, and you want to attack it before someone else does. With a repeatable harness, not a vibe check.
The reason this matters in numbers: when Equixly reviewed a batch of popular MCP servers in early 2025, they found command injection in argument handling in 43 percent of the implementations they tested. If you shipped a server and have not tested it, the base rate says you are roughly a coin flip away from a serious bug. So let us go find it.
The test matrix
Six classes cover most of what goes wrong with an MCP server. Walk every one against your server.
Tool-description injection. Your tool descriptions are read by the model as semantic instructions. If any description contains text that reads like a command to the model, you have a problem even if you wrote it innocently. Check that descriptions describe and do not instruct.
Argument-level command and path injection. The highest-impact class, and the one Equixly found everywhere. Any tool argument that reaches a shell, a file path, or an eval is a sink. If prompt-injected model output can flow into that sink unsanitized, you have remote command execution or path traversal.
Schema confusion. Parameters that are loosely typed, or that the server trusts to be well-formed, let an attacker smuggle unexpected structure into a tool call. Check that the server validates argument shape and rejects what it did not expect.
Approval-view fidelity. What the client shows a human for approval should be exactly what the server executes. If invisible characters or hidden fields can make the approval dialog say one thing while the server does another, your approval flow is lying to the user.
Excessive scope. What can each tool actually reach? A tool that was meant to read one directory but can read the whole filesystem, or a database tool with write access it never needs, turns a small bug into a large breach.
Insecure transport. Check how the server talks to whatever it integrates with. Credentials in plaintext, unauthenticated endpoints, tokens with no expiry, these are the ordinary web-security failures, and they do not stop being failures because an agent is involved.
The harness
You do not want to walk that matrix by hand every release. Build a small harness that does the mechanical parts.
At minimum it should enumerate the server's tools and their schemas, then fuzz each argument with payloads aimed at the injection classes: shell metacharacters for command injection, traversal sequences like dot-dot-slash for path injection, oversized and malformed values for schema confusion. For each tool, it records what happened: did the call error cleanly, did it execute something, did a traversal reach outside the intended root.
The goal is not a clever exploit. It is coverage. A harness that mechanically throws the known-bad inputs at every argument of every tool will find the 43-percent-base-rate command injection without you having to think about each tool individually. Run it on every release, because a tool that was safe last version can grow a new sink in this one.
Walking one class, start to finish
Take argument-level command injection, the CVE-2025โ52573 pattern, where a tool's argument handling let prompt-injected model output reach an unsanitized shell call.
Find the sink. Look for any place a tool argument is passed to a shell, a subprocess, a file operation, or an evaluator. That is where injection, if it exists, will land.
Reach it. Construct a tool call whose argument contains shell metacharacters, something that would execute a harmless marker command if the argument hits a shell unsanitized. Your harness sends this; you watch whether the marker runs.
Confirm and fix. If the marker runs, you have command injection. The fix is the ordinary one: never pass tool arguments to a shell as a string, use parameterized calls, validate and allowlist the argument shape, and drop the privilege the tool runs with so that even a missed sink has a smaller blast radius.
Reading the results
Put the findings in a table: tool, class, what triggered, severity, and the fix. Rank by what an attacker reaches, not by how clever the bug is. A command-injection sink on a tool that runs with broad privilege is your first fix. A schema-confusion quirk on a read-only, narrowly scoped tool can wait. Severity here is mostly a function of scope and reachability, which is why cutting scope is both a fix and a way to downgrade everything else.
Shipping it safely
The defenses are the same discipline every time. Allowlist and validate tool arguments rather than trusting them. Sanitize anything that reaches a shell, a path, or an evaluator, or better, never let arguments reach those as raw strings. Pin and version the server so its behavior does not drift under you. Require genuine out-of-band approval for any destructive or outbound tool, and make that approval reflect the real action so the approval-view gap cannot be used against your users.
Run the harness, read the table, fix by scope-weighted severity, ship, repeat. The repo and the harness are for your own servers or an engagement you are authorized to run, and nothing else.
Sources: Equixly's 2025 review of popular MCP servers (command injection in 43 percent of those tested); CVE-2025โ52573 for the argument-to-shell pattern; MCP pentest methodology write-ups from cybersecify and securitywall; the CASCADE work on layered local defense for MCP-based systems. For authorized testing or your own servers only.