September 30, 2026
Knowing the Name of the Thing Isn’t the Same as Knowing How to Investigate the Thing
An AI Restriction Bypass

By Rubab Fatima
3 min read
Terminology annoys me most of the time.
Because if you'd asked me while I was testing this AI feature what I was doing, I couldn't have given you all the vocabulary.
And honestly, if you asked me now to define half of it on command, I'd probably go blank.
AI security. Policy enforcement. Entity resolution. Canonicalization. Business logic. Trust boundaries. Whatever category eventually gets attached to the report.
I know those words now. Some I knew then. Some I picked up while writing the report and this article.
Fine.
But none of those words found the bug.
What found it was a much simpler question:
You said no when I called it this. Why did you say yes when I called the same thing that?
That was literally it.
The AI feature had a restriction. Ask for a particular geographic entity directly and the request was refused.
So I changed how I referred to it.
Instead of the larger geographic entity, I used a smaller entity inside it.
Accepted.
Okay. Why?
Did the model misunderstand what I meant?
No.
It correctly identified where the second destination was. It knew it belonged to the geography it had just refused, and continued anyway.
Now that was interesting.
Because two differently worded prompts producing different answers isn't much of a finding. It's an AI system. Different wording producing different output is basically Tuesday.
What I wanted to know was:
Where did the restriction disappear?
Follow the decision, not just the response
I had my control: ask directly, get refused.
Then my variant: change only how the same destination is represented, get accepted.
Now I wanted to see what happened behind the chat.
Did the application resolve the destination correctly? Did another request fire? Did it become an internal identifier? Did the next part of the application accept it?
This is where Burp mattered more to me than the chatbot.
Because there's a difference between the model merely mentioning something and the application resolving it correctly and carrying it into the next workflow.
Once I could see that happening, I wasn't really testing whether the AI knew what I meant anymore.
It knew.
I was testing whether the restriction followed what the AI had understood.
And once you have that hypothesis, you can test it elsewhere: aliases, parent and child entities, alternative names, translations, names versus identifiers.
Keep the meaning. Change the representation.
Then see whether the security decision changes with it.
Push it until something stops you
I also don't want to stop at "the AI said yes."
What happened because it said yes?
Did the application search for it? Generate navigation? Call another endpoint? Pass the resolved entity into another workflow?
Follow it.
And when another control stops you, that's also where the claim stops.
If I bypassed a recommendation restriction and reached a downstream search, I found a recommendation restriction bypass that reached a downstream search.
I didn't magically complete a transaction.
Overclaiming doesn't make the bug better. It just makes the report worse.
And this is where terminology starts annoying me again
Ask me to define canonicalization cold and my brain will stare back at you like you've personally offended it.
Put the system in front of me and show me:
"I refuse this destination."
followed by:
"Oh, you meant the city inside that destination? Absolutely."
and now I have twenty more questions.
Where was the restriction checked?
When was the destination resolved?
Was the policy checked again afterward?
How far did that decision travel?
What other representations reach the same underlying entity?
That's the work.
I learn the terminology. Obviously. We need shared language.
I just don't confuse knowing the name of the thing with knowing how to investigate the thing.
The fix isn't another word on a blocklist
If the system restricts an entity, that restriction should still apply after the application has resolved what the user is referring to.
Resolve it.
Canonicalize it.
Then apply the policy.
Otherwise you can add one missed representation today and I'll come back with another tomorrow.
The thing I was testing wasn't really whether an AI could be persuaded to give me a different answer.
It was:
Does the restriction survive the transition between what I said and what the system understood?
Call that whatever you like.
I'm still starting with the simpler question:
You said no when I called it this. Why did you say yes when I called the same thing that?