September 27, 2026
Building an XSS Scanner That I Can Actually Trust
Why reflection isnβt execution β and how context-aware testing can dramatically reduce false positives

By Arthur Johann Wilmsen Witt
8 min read
Why reflection isn't execution β and how context-aware testing can dramatically reduce false positives
If you have ever used an automated vulnerability scanner, you have probably seen something like this:
[HIGH] Cross-Site Scripting detected
Parameter: search
Payload: <script>alert(1)</script>[HIGH] Cross-Site Scripting detected
Parameter: search
Payload: <script>alert(1)</script>At first glance, this looks useful.
The scanner sent a payload, found something suspicious in the response, and reported an XSS vulnerability.
But there is a problem.
The payload appearing in the HTTP response does not necessarily mean JavaScript can actually execute.
And this is where a large number of XSS false positives begin.
While working on automated web vulnerability detection, I started thinking about a different approach:
Instead of asking, "Was my payload reflected?", the scanner should ask, "Where was my input reflected, how was it transformed, and can I actually break out of that context?"
That small change completely alters how an XSS scanner can be designed.
The naive approach
The simplest possible reflected XSS scanner looks something like this:
payload = "<script>alert(1)</script>"
response = requests.get(
target,
params={"search": payload}
)
if payload in response.text:
print("XSS detected")payload = "<script>alert(1)</script>"
response = requests.get(
target,
params={"search": payload}
)
if payload in response.text:
print("XSS detected")It works in vulnerable applications.
The problem is that it also reports many things that are not vulnerable.
Imagine the application returns:
<div>
You searched for:
<script>alert(1)</script>
</div><div>
You searched for:
<script>alert(1)</script>
</div>The input was reflected.
But it was HTML encoded.
There is no script execution.
The application could also return:
<textarea>
<script>alert(1)</script>
</textarea><textarea>
<script>alert(1)</script>
</textarea>or:
<script>
const search = "<script>alert(1)</script>";
</script><script>
const search = "<script>alert(1)</script>";
</script>or:
<input value="<script>alert(1)</script>"><input value="<script>alert(1)</script>">These cases may look similar to a string-matching scanner.
From a browser's perspective, however, they are completely different contexts.
That means the first principle of a better scanner should be:
Reflection is evidence. It is not confirmation.
Step 1 β Find reflection before testing exploitation
Instead of immediately sending XSS payloads, I prefer to start with a unique marker.
For example:
TESTING_38123TESTING_38123The scanner sends:
GET /search?q=TESTING_38123GET /search?q=TESTING_38123Then it searches for that exact marker inside the response.
If the marker never appears, aggressively testing that parameter for reflected XSS may not even be necessary.
If it does appear, we now have something much more useful:
a reflection point.
For example:
<p>Results for TESTING_38123</p><p>Results for TESTING_38123</p>or:
<input value="TESTING_38123"><input value="TESTING_38123">or:
<script>
let query = "TESTING_38123";
</script><script>
let query = "TESTING_38123";
</script>The next question becomes far more important than simply knowing the parameter is reflected.
What context are we inside?
Step 2 β Detect the reflection context
Context is one of the most important pieces of information in XSS detection.
Consider the following response:
<div>TESTING_38123</div><div>TESTING_38123</div>The user-controlled value is inside HTML text.
Now compare that to:
<input value="TESTING_38123"><input value="TESTING_38123">Here, the value is inside an HTML attribute.
And now:
<script>
const query = "TESTING_38123";
</script><script>
const query = "TESTING_38123";
</script>This time we are inside a JavaScript string.
Each situation requires a different breakout strategy.
A scanner should therefore classify reflections into contexts such as:
HTML_TEXT
HTML_ATTRIBUTE_DOUBLE_QUOTE
HTML_ATTRIBUTE_SINGLE_QUOTE
HTML_ATTRIBUTE_UNQUOTED
JAVASCRIPT_STRING
JAVASCRIPT_BLOCK
URL
JSON
CSS
HTML_COMMENTHTML_TEXT
HTML_ATTRIBUTE_DOUBLE_QUOTE
HTML_ATTRIBUTE_SINGLE_QUOTE
HTML_ATTRIBUTE_UNQUOTED
JAVASCRIPT_STRING
JAVASCRIPT_BLOCK
URL
JSON
CSS
HTML_COMMENTThis already eliminates one major design mistake:
using the same payload everywhere.
Why context changes everything
Take this payload:
"><img src=x onerror=alert(1)>"><img src=x onerror=alert(1)>If the application places input here:
<input value="USER_INPUT"><input value="USER_INPUT">the payload may be relevant because the first quote attempts to terminate the existing attribute.
But if the application places input here:
const search = "USER_INPUT";const search = "USER_INPUT";the same payload makes much less sense.
A JavaScript context might require reasoning about characters such as:
"
'
\
;
)"
'
\
;
)instead.
A good scanner should therefore select tests based on context rather than maintaining a giant list of payloads and throwing all of them at every parameter.
Step 3 β Character survival tests
Before attempting full payloads, the scanner can perform smaller tests.
Suppose the reflection is:
<input value="TESTING_38123"><input value="TESTING_38123">The important question becomes:
Can we escape the
valueattribute?
Instead of immediately sending an entire payload, we can test characters such as:
"
'
<
>
`"
'
<
>
`For example:
TESTING_38123"'><TESTING_38123"'><The server might return:
<input value="TESTING_38123"'><"><input value="TESTING_38123"'><">This tells us something extremely valuable.
The interesting characters are being encoded.
At this point the scanner can lower the confidence level without wasting dozens of requests trying equivalent payloads.
On another target, the result could instead be:
<input value="TESTING_38123"'><"><input value="TESTING_38123"'><">Now we have evidence that the surrounding syntax may be breakable.
That is much more interesting.
Step 4 β Understand transformations
Input is rarely returned exactly as it was submitted.
Applications may:
- HTML encode characters
- URL encode values
- escape quotes
- remove tags
- normalize whitespace
- remove keywords
- transform Unicode
- sanitize specific attributes
This means a scanner should compare the original probe with the reflected result.
A simplified model could look like:
sent = 'TESTING_"<>'
received = 'TESTING_"<>'sent = 'TESTING_"<>'
received = 'TESTING_"<>'From this difference, the scanner learns that:
" -> encoded
< -> encoded
> -> encoded" -> encoded
< -> encoded
> -> encodedThat information can be attached to the reflection.
For example:
{
"parameter": "q",
"context": "html_attribute_double_quote",
"reflection": true,
"transformations": {
"\"": "encoded",
"<": "encoded",
">": "encoded"
}
}{
"parameter": "q",
"context": "html_attribute_double_quote",
"reflection": true,
"transformations": {
"\"": "encoded",
"<": "encoded",
">": "encoded"
}
}Now payload generation becomes evidence-driven instead of random.
Step 5 β Context-aware payload generation
Once the scanner understands the reflection point, it can choose a very small number of relevant probes.
For HTML text:
USER_INPUTUSER_INPUTthe scanner might investigate whether HTML tags can be introduced.
For a quoted attribute:
<input value="USER_INPUT"><input value="USER_INPUT">the priority is testing whether the quote can be terminated.
For JavaScript:
let value = "USER_INPUT";let value = "USER_INPUT";the goal is understanding whether the JavaScript string can be escaped without producing invalid syntax before reaching a meaningful execution path.
The scanner is no longer asking:
Which one of my 300 payloads works?
It is asking:
What syntactic boundary is protecting this input?
That is a much better problem to solve.
Step 6 β Separate detection from confirmation
Another improvement is to stop treating vulnerability detection as binary.
Most scanners eventually return something similar to:
vulnerable = truevulnerable = trueor:
vulnerable = falsevulnerable = falseBut XSS testing usually produces a chain of evidence.
A better state model could be:
NOT_REFLECTED
REFLECTED
DANGEROUS_CONTEXT
BREAKOUT_POSSIBLE
PAYLOAD_INJECTED
EXECUTION_CONFIRMEDNOT_REFLECTED
REFLECTED
DANGEROUS_CONTEXT
BREAKOUT_POSSIBLE
PAYLOAD_INJECTED
EXECUTION_CONFIRMEDFor example:
REFLECTED
β
inside HTML attribute
β
quotes survive
β
attribute breakout succeeds
β
event handler injected
β
browser confirms executionREFLECTED
β
inside HTML attribute
β
quotes survive
β
attribute breakout succeeds
β
event handler injected
β
browser confirms executionOnly the final stages deserve high-confidence vulnerability reports.
Step 7 β Confidence scoring
Instead of marking everything as vulnerable, the scanner can calculate confidence based on evidence.
A simple scoring system could be:
Input reflected +1
Dangerous context +2
Dangerous characters survive +2
Context breakout detected +3
Executable structure injected +4
Browser execution confirmed +10Input reflected +1
Dangerous context +2
Dangerous characters survive +2
Context breakout detected +3
Executable structure injected +4
Browser execution confirmed +10A result could then look like:
Finding: Reflected XSS
Parameter: q
Context: HTML attribute
Reflection: confirmed
Quote breakout: confirmed
HTML injection: confirmed
JavaScript execution: confirmed
Confidence: 98%Finding: Reflected XSS
Parameter: q
Context: HTML attribute
Reflection: confirmed
Quote breakout: confirmed
HTML injection: confirmed
JavaScript execution: confirmed
Confidence: 98%Compare that to:
Parameter: q
Reflection: confirmed
Context: HTML text
HTML characters encoded
Execution: not observed
Confidence: 12%
Status: Reflection onlyParameter: q
Reflection: confirmed
Context: HTML text
HTML characters encoded
Execution: not observed
Confidence: 12%
Status: Reflection onlyBoth cases contain useful information.
Only one should become a vulnerability report.
Step 8 β Use a real browser for validation
Static HTTP analysis can go surprisingly far.
But eventually, if the goal is high confidence, a browser becomes extremely useful.
Tools such as Playwright can load the page and monitor runtime behavior.
For example, instead of relying on:
<img src=x onerror=alert(1)><img src=x onerror=alert(1)>the scanner could inject a controlled callback or marker and detect whether JavaScript execution actually occurred.
Conceptually:
window.__was_xss_confirmed = truewindow.__was_xss_confirmed = trueThe headless browser then checks:
window.__was_xss_confirmedwindow.__was_xss_confirmedIf it returns:
truetruethere is much stronger evidence of execution.
This changes the architecture significantly.
The scanner can perform fast HTTP-level testing first, and only send highly suspicious cases to browser validation.
Something like:
HTTP discovery
β
reflection detection
β
context analysis
β
breakout testing
β
candidate XSS
β
headless browser
β
confirmed XSSHTTP discovery
β
reflection detection
β
context analysis
β
breakout testing
β
candidate XSS
β
headless browser
β
confirmed XSSThat saves resources while keeping confidence high.
The DOM XSS problem
So far, most examples involve reflected server responses.
DOM-based XSS introduces another problem.
Consider:
const value = location.hash.substring(1);
document.querySelector("#content").innerHTML = value;const value = location.hash.substring(1);
document.querySelector("#content").innerHTML = value;If we visit:
https://example.com/#hellohttps://example.com/#hellothe server might never receive hello.
The vulnerability exists entirely inside the browser.
A scanner that only analyzes HTTP requests and responses might completely miss it.
This means serious XSS detection eventually needs some understanding of JavaScript sources and sinks.
Interesting sources include:
location.href
location.search
location.hash
document.URL
document.referrer
postMessagelocation.href
location.search
location.hash
document.URL
document.referrer
postMessagePotentially dangerous sinks include:
innerHTML
outerHTML
document.write
eval
setTimeout
setInterval
insertAdjacentHTMLinnerHTML
outerHTML
document.write
eval
setTimeout
setInterval
insertAdjacentHTMLBut even here, simple pattern matching is not enough.
Finding:
location.hashlocation.hashand:
innerHTMLinnerHTMLin the same JavaScript file does not prove data flows from one into the other.
Eventually, the interesting problem becomes data-flow analysis.
Reflections can occur multiple times
Another detail that is easy to underestimate: one parameter may appear in several places.
A request like:
?q=test?q=testmight result in:
<title>Search: test</title>
<input value="test">
<script>
analytics.search = "test";
</script><title>Search: test</title>
<input value="test">
<script>
analytics.search = "test";
</script>This means there are three separate reflection points:
HTML text
HTML attribute
JavaScript stringHTML text
HTML attribute
JavaScript stringThe scanner should analyze each reflection independently.
One may be completely safe.
Another may be exploitable.
Treating the entire response as one reflection can hide important differences.
Parsing HTML beats regex
A tempting solution is to identify context using regular expressions.
For simple pages, this can work.
But HTML is extremely messy.
Real-world pages contain:
- malformed markup
- embedded JavaScript
- templates
- inline JSON
- unusual attributes
- dynamic content
Using an actual parser provides a much stronger foundation.
The scanner can locate the unique marker in the DOM and determine:
parent element
attribute name
attribute quoting style
script context
surrounding charactersparent element
attribute name
attribute quoting style
script context
surrounding charactersInstead of trying to reconstruct all of HTML syntax using regex.
Regex can still be useful.
It just should not be responsible for understanding the entire document.
Stored XSS changes the architecture
Reflected XSS is relatively straightforward:
send request
β
receive response
β
look for reflectionsend request
β
receive response
β
look for reflectionStored XSS is different.
The injected data may appear several requests later.
For example:
POST /profile
bio=<payload>POST /profile
bio=<payload>The response contains nothing interesting.
But then:
GET /users/arthurGET /users/arthurcontains the stored value.
Testing this automatically requires the scanner to understand application state.
It must know:
where data is written
where that data can later appear
which user can see it
whether authentication is required
whether another role sees the resultwhere data is written
where that data can later appear
which user can see it
whether authentication is required
whether another role sees the resultAt this point, XSS detection starts moving toward application behavior modeling rather than simple fuzzing.
A practical scanner architecture
If I were building the system from scratch, I would separate it into stages.
Phase 1 β Parameter discovery
Collect:
query parameters
POST parameters
JSON fields
headers
cookies
form inputsquery parameters
POST parameters
JSON fields
headers
cookies
form inputsPhase 2 β Reflection probing
Send unique markers.
Record every reflection point.
Phase 3 β Context classification
Determine whether the marker appears in:
HTML
attribute
JavaScript
JSON
URL
CSS
DOMHTML
attribute
JavaScript
JSON
URL
CSS
DOMPhase 4 β Transformation analysis
Test characters and identify encoding or sanitization.
Phase 5 β Context-specific probes
Send only payloads that make sense for that location.
Phase 6 β Breakout detection
Determine whether the scanner actually escaped the original syntax.
Phase 7 β Browser validation
Use a headless browser for high-value candidates.
Phase 8 β Evidence generation
Save exactly why the scanner believes the target is vulnerable.
Evidence is more valuable than a severity label
One lesson I have learned while building security automation is that a scanner should explain itself.
Instead of:
HIGH β XSSHIGH β XSSI would rather receive:
Parameter:
search
Request:
GET /search?search=...
Reflection:
HTML attribute `value`
Encoding:
< and > preserved
" preserved
Breakout:
Confirmed
Injected node:
<img>
Event handler:
onerror
Browser validation:
Executed successfullyParameter:
search
Request:
GET /search?search=...
Reflection:
HTML attribute `value`
Encoding:
< and > preserved
" preserved
Breakout:
Confirmed
Injected node:
<img>
Event handler:
onerror
Browser validation:
Executed successfullyNow I can reproduce the finding.
I can understand why it was reported.
And, most importantly, I can trust the scanner more.
Fewer payloads can produce better results
It might sound strange, but improving a scanner can actually mean sending fewer requests.
Imagine two scanners.
Scanner A sends:
350 XSS payloads350 XSS payloadsagainst every parameter.
Scanner B sends:
1 reflection marker
1 transformation probe
2 context-aware payloads
1 browser validation1 reflection marker
1 transformation probe
2 context-aware payloads
1 browser validationScanner A has more payloads.
Scanner B has more information.
This is an important distinction.
Blind fuzzing optimizes for coverage through volume.
Context-aware scanning tries to maximize information gained per request.
For modern automated vulnerability detection, I think the second strategy is often far more interesting.
The goal is not to prove that input is strange
A scanner can detect strange application behavior very easily.
The real challenge is proving exploitability.
There is a massive difference between:
My input appeared in the response.My input appeared in the response.and:
I controlled browser syntax sufficiently to execute JavaScript.I controlled browser syntax sufficiently to execute JavaScript.A reliable scanner needs to understand that difference.
That means treating XSS detection less like string matching and more like a sequence of hypotheses:
Is the input reflected?
If yes, where?
What protects the surrounding syntax?
Can those protections be bypassed?
Can I escape the original context?
Can I create an executable construct?
Does the browser actually execute it?Is the input reflected?
If yes, where?
What protects the surrounding syntax?
Can those protections be bypassed?
Can I escape the original context?
Can I create an executable construct?
Does the browser actually execute it?Every step eliminates possibilities.
And that is exactly what makes the final result more trustworthy.
Final thoughts
When I first thought about automated XSS detection, the obvious approach was:
Send payloads and see which ones come back.
The more I worked with scanners, the less satisfying that approach became.
A better system does not start with exploitation.
It starts with observation.
It learns where input travels, how the application transforms it, what syntax surrounds it, and which boundaries can actually be crossed.
Only then does it attempt confirmation.
The objective should not be to send the largest payload list possible.
It should be to progressively eliminate hypotheses until the scanner has enough evidence to make a finding worth a human's attention.
Because in vulnerability scanning, finding more issues is useful.
But finding issues you can actually trust is much more valuable.