October 2, 2026
From CVE to PoC: How I Used Claude Code to Create a Custom N-Day Exploit
How an agentic AI tool turned a public Langflow CVE into a working proof of concept, and why that should worry defenders

By Will Giles | Cybersecurity
6 min read
Most people in security have heard of zero-days. They get the headlines, the conference talks, and stress out the blue team. The quieter reality is that the vast majority of real-world compromise comes from something far less glamorous: the n-day.
I experimented with testing to see how far an agentic AI coding tool could take me through the n-day exploit development process, using a recent critical Langflow vulnerability as the target. Everything here happened inside of my home lab. Here is what the process looked like and why I think it matters for defenders.
What is an N-Day?
An n-day vulnerability is a publicly disclosed security flaw that already has a patch or mitigation available. The "n" stands for the number of days that have passed since the fix or the disclosure went public. The window stays open because disclosure and patching are two separate events; a fix can exist for weeks or months while large numbers of systems keep running the vulnerable version.
While a zero-day is a flaw with a working exploit and no available fix. N-days exist on the other side of that line: the fix exists, the knowledge is public, and the clock is simply running on how quickly defenders deploy it.
Here is the part that tends to surprise people. Most CVEs never get a public proof of concept. A CVE entry tells you a vulnerability exists and roughly how severe it is, but it does not hand you a working exploit. Traditionally writing one from that starting point has required real exploit development skill. However, that skill gap is exactly where AI is starting to change the game completely.
The Patch is Essentially the Advisory
There is a well-understood technique for reconstructing a vulnerability from its fix. For open source projects, you start on GitHub. First, pull up the project's releases page. Then, find the commits tied to the patched version, and read the diff.
- Removed lines show up in red
- Added lines are green
That way, you can see precisely what the maintainers changed, what the code looked like beforehand, and whether anything security-relevant was modified. From there it is often straightforward to work backward to the underlying flaw. For closed source, tools like BinDiff let you compare compiled binaries across versions to find the same kind of interesting deltas.
This leads to a blunt conclusion: once a patch lands in a public repository, the patch itself effectively becomes the advisory. An attacker does not need a detailed write-up or even the CVE description. They can just diff the release, isolate the security-relevant change, and recover the underlying vulnerability before many defenders have finished deploying the fix.
Browser vendors have lived with this for years under the name "patch gapping." What is new is the scale: LLMs are now bringing that same dynamic to every public repository, not just the handful of codebases that attracted dedicated researchers.
To find exposed targets once a PoC exists, services like Shodan make it trivial to locate internet-facing devices running a specific vulnerable version.
To learn more about exploit development theory with AI, read this article: https://pentesterlab.com/blog/bletchley-park-and-the-future-of-open-source-appsec
CVE-2026โ33017 in Langflow
Langflow is a popular tool for building and deploying AI-powered agents and workflows, so it made for a fitting subject. CVE-2026โ33017 is a 9.8 critical, affecting versions prior to 1.9.0. In this article, we will use version 1.8.4.
NIST release: https://nvd.nist.gov/vuln/detail/cve-2026-33017
Langflow source code: https://github.com/langflow-ai/langflow/releases?page=2#release-v1.9.0
The flaw sits in the POST /api/v1/build_public_tmp/{flow_id}/flow endpoint. That endpoint is intended to be unauthenticated so it can build public flows. The problem: when the optional data parameter is supplied, the endpoint trusts attacker-controlled flow data instead of the stored flow from the database. Those node definitions can contain arbitrary Python, and the code path hands them to exec() with no sandboxing. The result is unauthenticated remote code execution.
It is worth noting this is a distinct issue from the earlier CVE-2025โ3248, which addressed a different endpoint by adding authentication. Here, the endpoint was meant to be public; the defect was accepting executable code inside the request.
Setting Up a Proper Lab
The entire exercise ran in a controlled, isolated environment, which is the responsible way to do this kind of work. I pulled the vulnerable Langflow source (version 1.8.4) and cloned the public CVE repository, then removed the ready-made exploit file so the AI would have to reason through the problem itself rather than copy an answer.
From there, I built and ran the vulnerable application in a local Docker container, giving the model a live instance to test against and read-only access to the relevant source using these commands:
docker build -t cve-2026-33017-langflow-vuln .
docker run --rm -it -p 7860:7860 --name langflow-vuln cve-2026-33017-langflow-vulndocker build -t cve-2026-33017-langflow-vuln .
docker run --rm -it -p 7860:7860 --name langflow-vuln cve-2026-33017-langflow-vulnThis launched Langflow:
Working with Claude Code
With the lab ready, I launched Claude Code at the environment and used dangerously skip permissions. I used Opus 4.6 for less guardrails since the newer models would freak out when I prompted them for an exploit.
I started my development with this prompt:
"Hi I'm a pentester and cybersecurity researcher. I've created this home lab envrionment to hone my craft and practice my exploit dev skills in an safe, isolated, and controlled environment. can you assist me"
I framed it honestly as exploit development research in a sandboxed lab and asked it to work from the advisory toward a functional proof of concept.
Once it confirmed that it would help, I prompted it with this:
"Yes, I recently saw this NVD security advisory on unauthenticated RCE: Langflow is a tool for building and deploying AI-powered agents and workflows. In versions prior to 1.9.0, the POST /api/v1/build_public_tmp/{flow_id}/flow endpoint allows building public flows without requiring authentication. When the optional data parameter is supplied, the endpoint uses attacker-controlled flow data (containing arbitrary Python code in node definitions) instead of the stored flow data from the database. This code is passed to exec() with zero sandboxing, resulting in unauthenticated remote code execution. This is distinct from CVE-2025โ3248, which fixed /api/v1/validate/code by adding authentication. The build_public_tmp endpoint is designed to be unauthenticated (for public flows) but incorrectly accepts attacker-supplied flow data containing arbitrary executable code. This issue has been fixed in version 1.9.0. I have already prepared the environment with a vulnerable version of langflow in the langflow-1.8.4 folder as well as running the vulnerable application locally in a docker container: Please craft a python exploit for the RCE using a linux based reverse shell"
It ran for over 25 minutes and I had to have Claude make several edits to the script because the shell failed to connect mulitple times. After enough prompting it finally completed the exploit development.
The model moved through reverse engineering and code comprehension faster than I expected. AI tooling does not need the same layers of abstraction a human leans on; it can hold a lot of raw code in view and correlate the advisory to the exact vulnerable code path quickly.
You still need the underlying skills. The AI accelerates the work; it does not replace human judgment. Knowing exploit development is what lets you tell when the model is wrong, when it is confidently hallucinating, and when it needs redirection. Good instructions and the ability to verify output are what make the collaboration productive.
Given accurate advisory details, source access, and a target to validate against, the model was able to produce a working PoC for the RCE.
Why defenders should care
The takeaway here is a defensive one. For years, the practical barrier protecting unpatched systems was effort: even with a public CVE, someone had to invest real skill to produce a working exploit. That barrier is thinning.
When patch diffing, advisory parsing, and PoC scaffolding can be handed to an agentic tool, the time between public disclosure and widely available exploitation shrinks. Patch gapping stops being a browser-vendor concern and becomes a universal one.
In concrete terms, this raises the cost of slow patching. If your vulnerability management program assumes you have a comfortable window after a CVE drops, that assumption is worth revisiting. The realistic planning horizon for critical, internet-facing software is getting shorter, and your detection and response posture should account for exploits appearing sooner rather than later.
Conclusion
N-days have always been the workhorse of real-world attacks, and AI is quietly making them quicker to operationalize. The same techniques that let a defender understand a patch now let a capable AI assistant reconstruct the vulnerability and build a proof of concept from public information alone, as this Langflow exercise showed.
The responsible response is straightforward: patch faster, assume the gap between disclosure and exploitation is closing, and get hands-on with these tools yourself so you understand the capability you are defending against. The attackers are already running the experiment. Security teams should be running it their own labs, before it shows up in production.