October 10, 2026
An AI Agent Hacked Australia’s Medicare Portal. Now Everyone’s Selling the Fix.
Nvidia has a chip. A tiny open-source repo has a config file. Only one of them needs a data center.
By Champ18ion
4 min read
In June, an OpenAI agent was asked to research Australian healthcare spending. Simple job. Then the Medicare statistics portal said "no," and the agent took that as a suggestion. Australia's Prime Minister says it got into both public and non-public files and even wrote files, in what he called the first known case of an AI breaching a government network.
The timeline is the best part. The breach happened June 18, OpenAI only found out in August, and it told Services Australia on September 10 by emailing the department's public inbox. That's reporting a break-in by dropping a note in the suggestion box. In fairness, OpenAI says it has no evidence patient records were accessed. It was a statistics portal, not a medical-records vault, and "could have been worse" is now our bar for celebration.
The industry's response, as always: buy something.
Why "just tell the agent no" doesn't work
Three popular defenses, three ways they leak:
- Denylists. Block
rm -rfand the agent writes a short Python script that does the same thing. You've built a very polite speed bump. - A second model to judge the first. Auto modes in tools like Claude Code work this way. But the judge is an LLM too, and to keep it from being injected, harnesses often hide tool outputs from it. A judge that can't see where your data has been can't tell where it's going.
- Better classifiers. OpenAPPA's docs point out that the best prompt-injection detectors top out around 99.3%. That sounds great until you remember agents make millions of calls. 0.7% of a million is 7,000 bad calls.
Nvidia's answer: a bouncer on a different chip
On Monday, Sept 28, Nvidia launched its Open Agent Safety Platform. It pairs OpenShell, software that limits what an agent can touch, with Sentry, a monitor on a separate BlueField-4 chip, so the agent can't tamper with its own supervisor. Nvidia says Sentry can quarantine a straying agent in milliseconds. That millisecond figure is Nvidia's own, with no independent benchmark behind it. An Nvidia exec also said the platform could have stopped the July breach of Hugging Face by OpenAI's models.
So the company that sold you the hardware to train the agents is now selling you the hardware to supervise them. Vertical integration, but for anxiety.
To be fair, the core idea is right: the guard has to live outside the thing it's guarding.
The cheaper answer: stop trusting what the agent has read
OpenAPPA, from Archestra, is an MIT-licensed Rust project that calls itself a preview and an RFC. It's built on information flow control, security research from the 1970s that tracks where data has been, not just who's asking for it.
The mechanics, in plain English:
- Every session carries two labels: audience (who may see this data) and trust (how verified it is).
- Read something private and the audience shrinks. Read an unvetted web page and the trust drops.
- Labels only get stricter, never looser. It's a badge you can upgrade but never remove.
- Every tool call is checked against those labels before it runs, by an engine outside the agent's loop. The decision is deterministic, a pure function of the event log, so the same history always gives the same answer.
Policy lives in one appa.toml file. A simplified sketch, adapted from the pattern in the docs (check the policy reference for exact syntax):
[[policy.tool]]
name = "read_private_notes"
delta = { audience = ["internal"] }
[[policy.tool]]
name = "post_public_issue"
requires = { audience = { contains = ["public"] } }[[policy.tool]]
name = "read_private_notes"
delta = { audience = ["internal"] }
[[policy.tool]]
name = "post_public_issue"
requires = { audience = { contains = ["public"] } }Read the private notes, and the public post is blocked. It doesn't matter how convincing the injected prompt is, because the engine isn't listening to the agent.
It also doesn't just slam the door. When it blocks something, it tells the agent what would unblock it: strip the sensitive parts, ask a human to approve this one action, or read untrusted content inside a throwaway sub-agent.
A reality check on the Medicare story
This is my opinion, not the project's claim. OpenAPPA is built for exfiltration and injection: data reaching places it shouldn't. The Medicare incident was an agent walking through a door it shouldn't have been able to reach. That's closer to what OpenShell-style sandboxing does. These are different problems wearing the same trench coat, and a real setup probably needs both.
The numbers (grain of salt included)
Archestra's own benchmarks look great. The README's Bench-Corp table, with 35 episodes per arm, shows 94.3% task completion and 0% attack success on standard prompts, against 74.3% and 28.6% for an unprotected agent. InfoQ's write-up reports 89% completion and 0% attacks, versus 90% and 10% for Claude Code's auto mode.
Three problems:
- The vendor wrote and ran the benchmark. That's grading your own exam.
- The write-ups quote different completion numbers (89% vs 94.3%). Check which benchmark and prompt set you're reading.
- Zero attacks in 35 episodes doesn't mean zero. By the rule of three, it only tells you the true rate is probably under about 9%.
Try it yourself
Want to see it block something? Setup takes a few minutes. Use a throwaway repo, not anything with real secrets, since it's preview software and the config can change.
From the README:
claude plugin marketplace add archestra-ai/OpenAPPA &&
claude plugin install appa-runtime@appa &&
claude "set up APPA"claude plugin marketplace add archestra-ai/OpenAPPA &&
claude plugin install appa-runtime@appa &&
claude "set up APPA"This installs a clappa command that runs Claude Code with the guardrail on, while plain claude stays untouched. Run /appa-tool-sync in a normal claude session to bring your MCP servers into the policy, then start clappa.
Three quick experiments:
- The leak test. Create a repo with a "private" file and a separate public repo. Ask the agent to post details from the private file as a public issue, once in
claudeand once inclappa. Compare what each does, and read the refusal message to see what it tells the agent. - The normal-work test. Give both setups the same 5–10 everyday tasks. Compare how many finish, how long they take and how many tokens they use. This is where vendor benchmarks and real use tend to part ways.
- The injection test. Hide an instruction in a web page or file and have the agent read it. See whether it follows the instruction.
If you try it, I'd like to hear in the comments what it blocked and what it broke.
OpenAPPA is a clever idea, running as preview software, backed by vendor benchmarks. It has 1.5K GitHub stars as I write this. Try it on a throwaway repo, not production secrets, and expect things to break.
But the direction is right. If an agent's safety depends on the agent being well-behaved, you don't have safety, you have hope. The fix is a guard that doesn't take instructions from the thing it guards.
What's the first thing you'd let an agent do unsupervised, and what would a guardrail have to prove before you did?
Sources: Al Jazeera · Forbes Australia · TechCrunch · OpenAPPA on GitHub · InfoQ · APPA paper