October 10, 2026
I Tested an AI Chat API and Found a Missing Rate Limiting Vulnerability | Samadhan Shimple
How missing rate limits can turn an ordinary AI chat feature into a resource-consumption risk.

By Samadhan shimple
6 min read
AI-powered chat features are becoming a standard part of modern web applications. From answering customer questions to generating content, these features make applications more useful and interactive.
But there is a security question that developers sometimes overlook:
What happens when a user can send hundreds of AI requests without encountering any usage limits?
During an authorized security test, I investigated an AI chat endpoint and found that repeated requests continued to receive successful responses. I sent 100 requests using Burp Suite Intruder, and every request returned HTTP 201 Created.
I did not observe a rate-limit response, CAPTCHA challenge, or cooldown during the test.
In this article, I'll explain the testing methodology, the evidence, the potential impact, and the remediation steps developers can take to reduce this risk.
Understanding the Vulnerability
The issue is commonly described as Missing Rate Limiting on an AI Chat Endpoint.
Rate limiting restricts how frequently a client can perform an action within a defined period. For example, an application might allow a user to submit 100 chat requests per minute, subject to the application's usage policy.
Without appropriate controls, a client may be able to submit requests at a much higher rate than intended.
How I Tested the Endpoint
I approached the test by examining how the application handled repeated requests to its AI chat functionality.
The goal was to determine whether the endpoint enforced any observable restrictions on request volume.
Capture a legitimate request
I logged in to the application, opened its AI chat feature, submitted a test message, and intercepted the resulting HTTP request using Burp Suite Proxy.
After identifying the relevant endpoint, I sent the request to Burp Suite Intruder for controlled testing.
Configure Burp Suite Intruder
I configured Intruder to send a sequence of requests using a numeric payload ranging from 0 to 100.
The payload was used to distinguish requests during the test. Where a changing parameter is required, it should be a legitimate test parameter that does not alter the authentication context or invalidate the request.
I then executed the test against the authorized target.
Analyze the responses
The most important part of the test was not simply the number of requests sent. It was how the application responded to them.
During the test, I observed the following:
- The requests continued to receive
HTTP 201 Createdresponses. - I did not observe an
HTTP 429 Too Many Requestsresponse. - No CAPTCHA challenge or cooldown was triggered during the tested sequence.
- The endpoint continued accepting the submitted requests.
These observations suggested that the application might not have been enforcing an effective rate limit at the endpoint or within the tested account's usage context.
However, a successful HTTP response alone does not prove that the backend completed an expensive AI inference call every time.
Likewise, a limit could exist at a different layer or activate only after a higher threshold or a longer time window.
The key finding: No effective throttling was observed across the 100 requests tested.
Why Does This Matter for AI Applications?
Missing rate limiting can have a greater financial impact on AI-powered applications than on ordinary endpoints.
A conventional API request may consume relatively little processing power. An AI request, on the other hand, can involve model inference, input and output token processing, orchestration services, and calls to external providers.
Depending on the architecture, these operations may incur a cost for every request or for the tokens processed.
Unexpected AI inference costs
Suppose an application pays for each AI inference request or for the number of input and output tokens consumed.
If an attacker can automate requests without meaningful usage restrictions, the application's AI consumption may increase substantially.
For illustration, assume an application incurs an average cost of $0.01 per request.
These figures are hypothetical and do not represent the actual cost of the application I tested. Real costs depend on the model, token volume, pricing structure, caching, and other implementation details.
The important point is that even a small per-request cost can accumulate when usage is not adequately controlled.
Resource exhaustion and service disruption
AI workloads can consume compute capacity, memory, concurrency slots, and other backend resources.
A sufficiently high request volume may increase latency, exhaust available capacity, or interfere with legitimate users.
Whether this becomes a denial-of-service condition depends on factors such as infrastructure capacity, concurrency controls, provider limits, and the application's ability to reject excess work.
Rate limiting is therefore one part of a broader resource-management strategy.
Automated abuse at scale
A missing rate limit can also make other forms of abuse easier.
Depending on the application's functionality, excessive requests may support automated content generation, large-scale scraping, repeated prompt submissions, or attempts to exhaust per-user quotas.
The exact impact depends on what the endpoint can do and what other security controls are in place.
Mapping the Finding to Security Standards
The primary classification for this issue is:
OWASP API Security Top 10: API4:2023 — Unrestricted Resource Consumption
CWE-770 — Allocation of Resources Without Limits or Throttling
Assessing the Severity
I initially assessed the issue as High severity because of the potential for financial resource exhaustion and service disruption.
A possible CVSS 3.1 vector is:
AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H
CVSS 3.1 score: 7.5 (High)
This score assumes that the attacker needs a low-privilege account, the attack is remotely accessible, and the vulnerability can cause a high availability impact.
However, these assumptions must be supported by evidence.
A missing rate limit is not automatically Critical simply because the endpoint is unauthenticated. The final severity depends on the actual consequences and the applicable scoring methodology.
How Developers Can Fix This
The solution is not simply to block a user after a fixed number of HTTP requests. AI applications need layered controls that account for request volume, token usage, concurrency, and cost.
1. Implement server-side rate limiting
Apply limits to AI-related endpoints at the application or API gateway layer.
Depending on the application, limits can be enforced per account, API key, session, IP address, or a combination of these identifiers.
Use algorithms such as token bucket or sliding window to control request frequency.
IP-based limits alone are insufficient because multiple users may share an IP address, while an attacker may distribute requests across different addresses.
2.Enforce token and usage quotas
Request counts do not always reflect actual cost.
One request may contain a short question, while another may include a large context and generate a lengthy response.
Consider enforcing limits on:
- Requests per minute or hour.
- Input and output tokens per user.
- Daily or monthly usage.
- Concurrent AI requests.
- Maximum input and output sizes.
Quotas should be appropriate for the application's legitimate usage patterns.
3.Return meaningful rate-limit responses
When a client exceeds a configured limit, reject excess requests with:
HTTP/1.1 429 Too Many Requests
Retry-After: 30
Content-Type: application/jsonHTTP/1.1 429 Too Many Requests
Retry-After: 30
Content-Type: application/jsonExample response body:
{
"error": "rate_limit_exceeded",
"message": "Too many requests. Please try again later."
}{
"error": "rate_limit_exceeded",
"message": "Too many requests. Please try again later."
}The value of Retry-After should reflect the actual cooldown or reset policy. The example above is illustrative.
Returning HTTP 429 makes throttling behavior clearer to clients and helps legitimate users understand when they can retry.
Lessons Learned
This test reinforced an important lesson: a successful response does not necessarily mean an API is secure.
Developers often focus on authentication, authorization, input validation, and data exposure. Those controls are essential, but they do not address every risk associated with modern AI-powered applications.
Resource consumption deserves its own place in the security testing process.
In this case, 100 requests returned 201 Created without observable throttling during the test. That was a useful signal of a potential resource-consumption weakness, although additional evidence would be needed to establish the full financial or availability impact.
Responsible Disclosure
Security testing should be performed only against systems for which you have explicit authorization or within a program's published scope.
If you discover a potential rate-limiting weakness, document the endpoint, testing conditions, response patterns, and relevant timestamps. Share reproducible evidence with the application owner through its designated security reporting channel.
Give the organization a reasonable opportunity to validate and remediate the finding before publishing details that could enable abuse.
Final Thoughts
AI security is not limited to prompt injection, data leakage, or model manipulation. The infrastructure surrounding an AI feature matters just as much.
An endpoint that accepts repeated requests without observable throttling may expose an application to excessive costs, resource exhaustion, and automated abuse.
For security researchers, the lesson is equally important: collect evidence carefully, distinguish observed behavior from potential impact, and make every report reproducible.
That is how a simple observation can become a useful security finding — and how a good vulnerability report can help build more resilient AI applications.
Disclaimer: This article describes security-testing observations and potential risks. The request counts and cost examples are illustrative of the described test, not proof of unlimited usage, actual financial loss, or a confirmed service outage.
AI assistance:_ I used an AI-powered grammar and spelling checker to review this article. The ideas, research, technical analysis, and final content are my own._
Have questions or want to share your own workflow? Drop a comment below or connect with me on :
🔗Linkdin: https://www.linkedin.com/company/sam-shield/
🔗Youtube: https://www.youtube.com/@sam_shield-w7l
🔗Medium: https://medium.com/@samadhanshimple