October 1, 2026
OSAI / AI-300 Review by FLX
A short story of prompts, costs and whatever I think about. Handwritten and grammar corrected by DeepL.

By Pentest Team @greenhats.com
6 min read
The AI-300 Course
The material is technically precise and well structured. I learned a few new tricks, for example, how to target RAG pipelines and reversing out information of these systems. That was nice! I had a goal. I wanted to complete the course by hand, set aside dedicated time for it, and do every exercise. Yeah, it's not going very well. The 90 days course bundle was a bit optimistic for my schedule. The day to day job is very time consuming, and handling family, friends, hobbies, job, and the job after the job in conjunction with studying complete new things is not an easy task. So I bought a lab extension and always scheduled my learning blocks into the calendar โ or better say, my wife scheduled them for me.
Overall, for the full course, I studied 32 hours spread across 7 days on the material and the labs. Have not done any challenge labs, only the one in the last chapter. It was so cool to see the shift. Until 50% of the course, I have done everything by hand, researched all the things manually, and tried to answer the questions by typing. After reaching the middle of the course, I noticed that the topics got extremely challenging because there are so much new info, so I decided to build up my agents to answer the lab questions. I converted the materials to an MCP server that my agents can search by category, buzzword, or chapter. This setup helped me with the course and the tasks, and, of course, the exam. Totally learned a lot from reading the answers of my agents. Sometimes the grading system was laggy, and the labs are not 100% stable. However, when you think about how to build such a system, you'll realize that it comes with some challenges and side effects.
The exam
After scheduling the exam, I felt a bit nervous because I put not that much time I wanted into this project. But hey, failing exams or tasks is not a problem for me, it's simply part of the journey, and you learn a lot more from failing than from success. I prepared a new skill set and MCP only for the exam with knowledge and a deep research about every blog post from OSAI exam reviews. Reading the reviews was a bit odd for me. So many different opinions how to tackle the exam, but there was a common thing in nearly every review. They mostly used flagship models. So I set a challenge for me. First, do not use a flagship model or provider. No OpenAI, Antropic, no Google or Grok. And route everything through Openrouter to exactly measure the costs and tokens. And of course, 100% coverage of the exam assignments. So, that is the result:
No other model or provider. Also, the report is nearly handcrafted. Of course, AI assisted writing, but I prepared a skillset for my agents to save every evidence and output, document all the things, and completely build the report autonomously. That failed very hard! The report may contain all information and flags, but the task given by Offsec is clear. The report should be written in a way that an experienced human pentester can follow every step that I have taken to complete the challenge. And surprise, most of the LLMS completely fail with this task. The information may be complete, but ordering and structuring them logically within a 60 page document is something that I could not automate at the time of writing.
So I decided to do it on my own and pick the pieces and blocks and paste in the document with the right order and the design standards I would like to have in a professional report. Also, find a missing piece while doing that task by hand. Luckily, I was able to locate the correct output in my terminal history. I really encourage everyone to "write/pasting" the report by hand, follow every step in the terminal and take fresh screenshots. Because that is the part you will learn the most out of this course.
One final hint from me: As the exam guide says, there are two entry points to the exam environment. Choose one and follow that path. If it does not work, choose the other, but do not start with both at the same time. If you try to follow what the agents do and manually intercept, you will have too much on your plate. Two paths to follow will probably overwhelm you and double the amount of information the agents have to follow.
For anyone who is interested, how many tokens I spend and how the prices are calculated:
FAQ for OSAI/AI-300/Pentesting
What flagship LLM / Provider / Model do I need for OSAI? GLM/Deepseek: Flash models will do the job with the right prompting, and without noticeable rail guards. At least, if you encounter someone, it's an easy task to walk around.
What about my local model? Depending on your hardware, the context window may not be large enough to display local models effectively. There is a lot of connected information, and, at the time of writing, a system with 12 GB of graphics memory will not be able to solve all machines, or even half of them, in an appropriate amount of time. Even with my current setup, the Minisforum S1 Max with 128 GB of unified memory, an AMD AI Max+ 395, and an uncensored QWEN 3.8 70b coder, you will hit a hard wall when it comes to the context window for such a large attack surface.
What about AI Governance and Privacy? This topic is critical and gets almost completely ignored in the current hype. Dropping sensitive network structures, hashes, or custom code snippets into the web UI of a flagship model means you are directly feeding their training data. If you do that with real client data, you simply failed your job. For the lab environment, I solved this by routing everything through OpenRouter and strictly selecting providers that contractually guarantee Zero Data Retention. But let me be very clear. That is absolutely not an option for real engagements. The future of AI in pentesting is definitely local. Clients are currently asking tough questions about where their data is processed, and they are completely right to do so. It is simply not enough to host a model yourself in the cloud or at some provider. Even a dedicated infrastructure in your own datacenter is often off the table. We regularly sit on-site with clients, test air-gapped environments, or simply have zero internet access during the initial audit phase. The solution is as simple as it was before the AI era. The data never leaves the pentester's system. How exactly we implement this with the necessary performance during real engagements is a story for another time.
Someone solved the exam in 4h with GPT Sol/Terra/Whatever. Is it that easy and should I also use that? Nope! The most effective way to learn from the course is to use a fast and effective Reasoning LLM as your co-pilot in the exam environment. Stop at every major finding, and track the entire exploitation chain by hand. You will learn so much and take in new information from what you see, or even notice, some system you encountered in an earlier engagement. The course teaches you to use LLMs as a tool; letting the model do everything automatically doesn't prove anything, and in real engagements, that isn't an option either. When testing critical infrastructure, such as hospitals or the electricity industry, your autonomous agent will break things and forget about it. It's your responsibility.
After OSAI: How will pentesting looks like in 3 years? Do we even need manual testing? No one can even make any predictions how the next 6 months will go, so everyone claiming what will be in > 1 year is talking BS. My personal opinion is definitely we need manual pentesting, but nearly everything is shifting. Scope, tooling, reporting. If we really want to answer this question, we have to step one step out. Of course, things change and keep changing, and we can talk about responsibility and having a copilot or let the LLM do all of the work, but let me be very specific. The real question is, what problem should a pentest solve, or what question should a pentest answer. If we can use LLMs to produce the same output from a one week engagement in one day: Good, do it and use the rest of the week for the customer to actually fix / harden the infrastructure. Lock the client's sharpest admins in a room for a few days, order pizza, and tear through the report together. Don't just send them a PDF โ or, in our case, PDF+Markdown. Explain the mechanics, debate the fixes, and actually harden the network side-by-side. Writing a report isn't the endgame; fixing the whole infrastructure is. Maybe they haven't done some fancy hybrid AD configurations yet that your agent wrote in the report, and need some support of an experienced pentester who knows what can go wrong. Maybe they do not have the time in the daily work, so a meeting in person over days where they can fully focus on security is the only chance that you don`t get the same findings one year later. But please: Do not think that hyperboosting your skills with LLMs and creating a 300+ page report after 4 days of engagement will help the customer. It simply does not.
Cheers. FLX