October 1, 2026
How I Use AI to Understand Applications Instead of Just Generating Payloads
At the onset of using AI for bug bounties, I asked the same questions that everyone else does. “Please give me XSS payloads in 50 samples.”…

By Sairaj Thorat
5 min read
At the onset of using AI for bug bounties, I asked the same questions that everyone else does. "Please give me XSS payloads in 50 samples." "Please give me some SQL injection strings for this particular parameter." "What do you suggest testing on this login form?"
Productive though it seemed, all my notes were filled with a bunch of payloads, and no longer did I have the problem of facing an empty page. However, after a couple of weeks, it hit me that I was now armed with a ton of things to test while learning absolutely nothing new about my targets. The AI was there to help me test my hypotheses but not to guide me on whether my hypothesis was right.
Thus, I modified my question. "Please tell me how is this application functioning?".
Why Payload Lists Fall Short
The payload is the final stage of the vulnerability, not the starting one. Prior to applying a payload, many things need to be in place — there needs to be an input reaching the sink, the application must process the input in a certain way, and none of the actions in-between must nullify it.
Creating a payload list completely disregards all this. A payload list presupposes that the payload itself is what's interesting in the whole attack scenario, while in fact it is the behavior of the application which is what's interesting, and a payload list doesn't tell you anything about it at all.
Secondly, using payload as a starting point is noisy. You send a bunch of requests, get an identical error message, and it becomes unclear whether your target is secure or whether you've just hit the wrong spot. A non-result tells you nothing because you don't even know what the right result was supposed to be.
Understand First, Test Second
Before I get to a new feature, I would like to know a few things:
. What does this feature do, and what endpoints does it use? . What are the user-controlled parameters, and which ones are automatically set by the client? . Where will the provided data be sent, and who will be able to access it? . What assumptions are being made regarding input to the application? . What checks take place in the browser, and which need to happen on the server?
A large JavaScript file is the ideal venue for posing these questions. Modern front ends contain much logic — builders for requests, validation rules, feature toggles, secret routes, and permissions checks. Browsing the entire thing manually takes hours, but an AI can go through it in minutes.
The magic happens in the way I phrase the question. Instead of simply throwing the file at the AI and asking it to "hunt for vulnerabilities," I start by asking it for the narrative — the purpose of each function, the way functions relate to one another, and how data flows from the input field to the server.
Scenario One: Fields the User Never Typed
Here is an example of what might be found in the JavaScript of a profile page. It is a made-up scenario just for demonstration:
function saveProfile(form) {
const body = {
name: form.name.value,
bio: form.bio.value,
userId: currentUser.id,
role: currentUser.role
};
fetch("/api/v2/profile/update", {
method: "POST",
body: JSON.stringify(body)
});
}function saveProfile(form) {
const body = {
name: form.name.value,
bio: form.bio.value,
userId: currentUser.id,
role: currentUser.role
};
fetch("/api/v2/profile/update", {
method: "POST",
body: JSON.stringify(body)
});
}The weak prompt: "Tell me the payloads to attack this endpoint."
The improved prompt: "Tell me what this function sends to the server. What fields are user-supplied, and what are client-generated? What checks are required in the backend for this to be a secure request?"
An effective response will note that userId and role are sent in the request despite not being entered by the user. It shows an assumption in the design in which the client tells the server its identity and permissions and the server takes it on trust.
Here are my testable questions instead of a list of payloads:
. Does the server check for ownership if I change the userId to one of another user's accounts? (This could be IDOR.) . How does the server react when I change the role? (This could be privilege escalation.)
The tests happen in Burp. Either everything is validated, and I'm done. Or it isn't, and I'm on to something. Whatever, I now know what that feature does, and there was a purpose behind each of my requests.
A Small Prompt Toolkit:
These are the prompts I reuse most. Adjust them to your target:
. Feature summary: "Summarize what this code does in plain language, then list every network request it can make."
. Trust analysis: "Which values in these requests are controlled by the user, and which are set by the client? Which would the server need to verify?"
. Data flow: "Trace this parameter from where it enters to where it's used. Flag any place it's validated, encoded, or transformed."
. Assumption hunting: "What assumptions does this code make about its input, and what would happen if each assumption were false?"
. Inconsistency check: "Do any two code paths handle the same input differently?"
Notice that none of these ask for an exploit. They ask for understanding, and understanding is what makes the next move obvious.
Treat Every Answer as a Hypothesis
However, there are some limitations, and overlooking them will be a waste of your time at best.
It invents things. AI can describe endpoints, parameters, or flaws that aren't in the code, and it does so confidently. Confirm everything against real traffic before believing it. It can misread logic. An explanation that sounds right may misunderstand control flow, especially in dense or heavily abstracted code. Big or minified bundles break it. Large files may exceed what the model can handle. Split them into chunks, and be skeptical of anything it says about obfuscated code. Confidentiality is real. Pasting a target's code into a third-party service may break program rules or expose sensitive data. Read the policy first and strip identifying details.
So I consider all explanations as hypotheses. Finding an apparent vulnerability and proving its existence are two entirely different things, and it is my job to prove it.
Common Mistakes
A few traps I've fallen into, so you don't have to:
Believing the summary instead of reading the code. If a finding looks promising, I go read the actual lines myself. The AI points; I verify. Asking one giant question. "Analyze this whole app" produces vague output. Small, specific questions about one feature or one parameter produce answers you can act on. Skipping the manual proof. An AI saying "this looks like an IDOR" is not a report. A reproduced request and response is. Chasing every lead. Understanding a feature produces many questions. Pick the ones with the highest impact and the weakest-looking assumptions first.
Where Payloads Still Make Sense
That being said, this does not mean that payload generation has no use at all. When I am aware of where the inputs go and how they are handled, it is precisely then that a payload would be perfect. I can request payloads that work when the inputs are used within a JavaScript string, as well as those that survive certain filters.
The distinction here lies in order. First comes the understanding, which means there is an intended use for the payload.
The Takeaway
It is not about getting AI to think for me. It is about making it possible to skip over the time it takes to analyze and understand the application itself and to focus on what is truly important — the validation of the assumptions that seem weakest.