October 1, 2026
Fine. Call It AI Security. Iโm Still Testing the Trust Boundary.
What Changed When an LLM Got Put in the Middle of the Web Application

By Rubab Fatima
6 min read
Apparently putting an LLM in the middle of a web application means I now need an entirely new vocabulary.
Prompt injection. Indirect prompt injection. Excessive agency. Improper output handling.
The terminology is useful.
But after actually playing with these attack paths, I keep coming back to a much less exciting diagram:
User โ AI feature โ LLM โ Application/API โ Data
And then I start asking the same annoying questions I was asking before somebody put an LLM in there.
What can I control? What can this thing reach? What authority does it have? What trusts its output? And where is the actual security decision being made?
The chatbot is probably not the thing you're actually testing
This is the part I think gets muddled surprisingly often.
A lot of companies didn't build the underlying LLM. They bought access to one.
They take a third-party model, connect it to their application, give it context, feed company data into it, expose functions or tools to it, connect those to APIs, and put a nice chat interface in front of the whole thing.
So what I'm looking at might be closer to:
Me โ Company's AI assistant โ Third-party LLM โ Company API โ Company data
And suddenly saying "I'm attacking the LLM" becomes a little misleading.
If I manipulate Company X's AI assistant into doing something interesting with Company X's data, that doesn't automatically mean I've found a vulnerability in whichever company provides the underlying model.
I'm testing Company X's integration of that model into its application.
That's where the context was added, the tools were exposed, the data was connected, the permissions were granted, and somebody decided what happens when the model asks for something.
The underlying LLM might be third-party.
The trust you gave it isn't.
And that's where the wrapper becomes interesting
"It's just an AI wrapper" gets used almost dismissively.
From a security perspective, just a wrapper can contain an impressive number of decisions I want to poke.
What did you put in the system context? What company data can the model see? Which functions can it call? What credentials are those calls made with?
And if I can influence something the model sends to a function, and that function eventually sends my input to an API, I'm going to follow it.
Can I reference something belonging to another user? Can I reach functionality that isn't directly exposed to me? Can I pass something through the model that becomes dangerous when another component processes it?
At that point, I don't really care that there's an LLM sitting in the middle.
I'm following the input.
Because sometimes the LLM hasn't created an entirely new vulnerability.
It has created a new route to an old one.
Put path traversal behind an LLM and path traversal doesn't suddenly become AI traversal. Put a vulnerable API behind a chatbot and the API doesn't stop needing access control. Put an internal service behind a tool call and I'm still interested in what I can make that service reach.
The architecture changed.
The security questions didn't disappear with it.
Who in the vibe-coded world gave this thing permission to do all that?
This is basically how my brain translates excessive agency.
If the assistant needs two functions, why can it call seven?
If it can read customer information, whose information?
If it performs an action, under whose authority?
Because now we're right back in one of my favorite rabbit holes:
Identity. Permissions. Authority. Trust.
Something is acting on behalf of something else. I want to know whose authority travelled with that action, and whether the component that actually owns the resource checks that authority for itself.
And this is where the confused deputy problem starts making sense to me.
Something has authority I don't have, and I can influence it into using that authority for me.
Except now that something is an LLM connected to tools, and may decide what I meant before making the call.
Which brings me back to one question:
If I ask it to do something I'm not authorized to do, who says no?
The model?
The API?
Hopefully not just the model.
The model isn't your authorization layer
Suppose I shouldn't be allowed to access another customer's information.
You can tell the LLM:
Never provide one customer with another customer's data.
Lovely.
Now I'm going to assume I get around that.
Not because I know that I can.
Because that's the security question.
What happens if I do?
The API holding that customer data should still know that I am me, the resource belongs to somebody else, and I am not authorized to access it.
The model can decide to comply with my request.
That still shouldn't mean the API lets me access something I'm not authorized to access.
The model can be another control. But it shouldn't be the control protecting somebody else's data from me.
The authorization still has to happen where the data or action is actually controlled.
That's access control.
The LLM just gave us another place to accidentally forget it.
And that's also why I'm much less interested in making the LLM impossible to manipulate than I am in:
If I manipulate it successfully, what did you allow me to reach through it?
That's where the impact lives.
Now flip the trust around
So far I've been looking at:
User โ LLM โ Company systems
But an LLM doesn't only consume what I type into the chat box. It may consume API responses, documents, emails, web pages, search results, knowledge bases, other user-generated content.
And now the question becomes:
What does the model trust?
This is where indirect prompt injection gets properly weird.
I don't even have to manipulate the model directly. I can control something it reads later.
So now the path looks more like:
Attacker-controlled content โ LLM โ privileged function
And now I have attacker-controlled content influencing something that can actually perform privileged actions.
Because something that one part of the application considers data may be interpreted by the model as instructions.
And this, to me, is one of the bits that is actually different.
The application can know that one piece of text came from the user, another came from a retrieved document, another came from somewhere else entirely. The model still has to interpret what all of that means and what, if anything, it should do about it.
There's no parameterized query for please know which sentence is trying to manipulate you.
So now I'm not only asking whether attacker-controlled data can reach something sensitive.
I'm asking:
Can attacker-controlled data change what the thing with access thinks it's supposed to do?
That's why prompt injection isn't SQL injection with a trendy name.
With indirect prompt injection, the attacker might never interact with the model at all. They only need control over something the model later consumes.
And for me, the important question is what can the model actually do after that content influences it? If all I can change is what the chatbot says, that's one problem. If the model has access to privileged functions, that's a very different problem.
Then there's the other side of it: what does the application do with whatever the model produces?
Does it display it on a page? Pass it into another function? Use it to perform an action?
Because if I can influence the model's response, I want to know whether something else in the application blindly trusts that response.
The LLM doesn't magically make my influence safe just because it passed through the model.
And this is where improper output handling comes in.
This doesn't mean AI security is just web security with extra steps
It isn't.
There are genuinely different problems here.
Models are probabilistic, and that makes the trust boundary weird.
The questions are familiar.
The answers aren't always deterministic anymore.
Is this data or an instruction? Which function does this sentence mean I want? What arguments should be generated? What happens when the thing making those decisions interprets something differently than I expected?
Giving models tools and agency changes what successful manipulation can actually accomplish. Training data, retrieval systems, model behavior and other parts of the AI stack introduce attack surfaces that aren't covered by drawing a few API boxes.
There is new stuff to learn.
What I don't buy is the idea that putting an LLM into an application somehow requires me to throw away everything I already know about investigating systems and start again.
No.
Show me the new component.
I'll learn how it behaves.
Then show me what you connected it to.
I don't need the LLM to behave
I think this is the assumption I'd rather start with:
Eventually, the model will do something you didn't want it to do.
Maybe somebody manipulated it. Maybe external content influenced it. Maybe the model behaved unpredictably. Maybe your instructions weren't as clever as you thought.
Now what?
Does the API still enforce authorization? Does the application validate the action? Does the model have access only to what it needs? Does a sensitive action require something stronger than the LLM deciding that my sentence sounded convincing?
Because that's the bit I actually care about.
The model can be probabilistic.
The security decision can't afford to be.
I'm still following the trust
I'll test prompt injection. I'll map the functions and follow the APIs. I'll see what happens when attacker-controlled content gets into the model's context, and what happens to its output when something else trusts it.
And yes, I'll learn the terminology for whatever I find.
But I'm still going to draw boxes and arrows, ask who controls what, follow authority across components, and poke the place where one thing assumes another thing is safe.
Because companies can outsource the model. They can buy the API. They can plug somebody else's intelligence into their application.
What they cannot outsource is the security of the trust they build around it.
Put an LLM in the middle of the web application if you want.
You just gave me another trust boundary to test.