What We’ve Learned Building CASM-AI: An AI Pentesting Platform

Oct 2, 2026

Author: Brian Judd

CISSP, CISA, CRESC - VP Information Assurance

I’ve spent a large part of my career around penetration testing, and one thing has always been true: there is never enough time to test everything as deeply as we would like. A good pentester can do a lot in a week, and a great pentester can notice the one strange response, unexpected service, odd authentication behavior, or small configuration mistake that turns into a meaningful attack path.

But even the best tester still has only so many hours. That limitation is what made AI interesting to me in the first place.

Over roughly the past nine months, I’ve had the opportunity to play a major role in SynerComm’s research and development around AI-assisted penetration testing. We started with a fairly simple question:

How much of the penetration-testing process can AI meaningfully perform if we give it real tools, a methodology, controlled access to targets, and enough freedom to investigate what it finds?

The answer has surprised me. That work has become CASM-AI, our agentic penetration testing platform.

I want to be careful with that description, because “AI pentesting” can mean almost anything right now. We are not talking about asking a chatbot to explain a vulnerability, using AI to rewrite a report, or replacing experienced penetration testers. We are talking about AI agents that can actively participate in the testing process: enumerate systems and applications, use offensive-security tools, analyze results, develop hypotheses, perform additional testing, collect evidence, and keep investigating.

That distinction matters.

 

The Part That Impressed Me Most

Running tools is easy. We have been automating Nmap, vulnerability scanners, web crawlers, content discovery tools, and countless other security utilities for decades. What makes the current generation of AI interesting is what happens between the tool executions.

A traditional scanner runs a check and compares the response to something it already knows. An AI agent can look at a result and ask a new question:

  • Why did that endpoint respond differently?
  • Why does this unauthenticated request return nearly the same object as the authenticated one?
  • Why is this application referencing another hostname?
  • Is that JavaScript file exposing something useful?
  • Does this service actually require authentication?
  • Is the behavior I just observed consistent across users, objects, or requests?

That reasoning loop is much closer to how a pentester works: observe something, develop a hypothesis, test it, evaluate the result, adjust, and try again.

We have already seen agents follow exactly that kind of path in practice. In one case, an agent noticed that an unauthenticated request was behaving almost like an authenticated one. Instead of simply recording the odd response, it kept testing the access boundary until it had enough evidence to surface a real access-control issue for human validation.

The AI is not just running a list of checks. It can get curious.

That does not make it a senior penetration tester. It does, however, make it far more useful than the traditional automation we have had available to us, and it can repeat that process relentlessly.

 

We Are Not Trying to Replace Pentesters

Our objective is not to reduce the need for penetration testers. It is to make our penetration testers more capable.

Experienced pentesters are expensive for a reason. The value is not in their ability to type an Nmap command or click through hundreds of web pages. The value comes from years of accumulated judgment: understanding what matters, recognizing when something does not make sense, knowing how far to push an attack, combining weaknesses, evaluating impact, and explaining risk to a client. Those are exactly the hours we want to protect.

If an AI agent can spend hours enumerating application functionality, analyzing large amounts of JavaScript, comparing responses, reviewing exposed services, chasing down suspicious behavior, organizing evidence, and eliminating dead ends, then our human tester can start from a much better position. Instead of spending valuable time asking what should I look at?, the tester can increasingly focus on which of these promising issues deserves my attention first?

The same applies to traditional vulnerability-scanner results. Aletheia, one of our specialized agents, is focused specifically on reviewing scanner findings and dispositioning them as likely real, likely false positive, or requiring additional human or agent validation before they consume a pentester’s time.

For our Continuous Penetration Testing clients, this becomes even more interesting. Their environments are constantly changing: new hosts appear, applications change, DNS records move, new endpoints are deployed, and authentication behavior changes. Traditional pentesting is inherently constrained by time. AI gives us a way to investigate more of those changes without pretending that every automated observation should become a finding.

 

Specialized Agents, With Humans in the Middle

We designed CASM-AI around specialized agents rather than trying to create one giant AI “super pentester.” We kept the Greek mythology names because, frankly, pentesters still like to have some fun with their tools.

  • Zeus handles orchestration and helps determine what needs to be tested.
  • Hermes focuses on web applications and APIs.
  • Argus focuses on networks, systems, services, and infrastructure.
  • Aletheia reviews vulnerability-scanner results and helps determine whether they appear valid, are likely false positives, or require further testing.
  • Other specialized agents support parts of the finding, reporting, and quality-assurance process.

Our pentesters interact with the platform through Olympus, our internal operations console. Olympus is proprietary, intended only for SynerComm pentesters, and is not exposed as a public Internet-facing portal.

The names are fun. The separation of responsibilities is not. Narrowing what each agent is responsible for makes the system easier to control, easier to monitor, and easier for a human tester to understand.

 

The Safety Problem Is More Interesting Than the AI Problem

One of the first lessons in this project was that telling an AI agent to “stay in scope” is not a security control. A prompt is an instruction, and instructions can be misunderstood. The controls that matter therefore have to exist outside the model.

We have approached CASM-AI with the same kind of defense-in-depth mindset we would recommend to a client. We do not want any single safeguard to be the only thing standing between an autonomous agent and an out-of-scope system. Our controls include multiple layers such as:

  • Hard-coded and programmatically enforced testing scope
  • Network access-control lists
  • Firewall rules that restrict what agent systems can reach
  • Separation of capabilities between agents
  • Logging of LLM requests and responses
  • Logging of commands issued by the agents
  • Logging of command output and execution results
  • Logging and monitoring of firewall activity
  • Retention of evidence and testing artifacts for review

During development, training, and controlled testing, we have also used full packet capture so we can see exactly what network activity occurred. Broader PCAP collection remains something we can use where it makes operational and security sense as the platform matures.

There is also an important distinction between autonomy and authority. Safe, non-destructive testing can run autonomously within the approved scope. Actions that are intrusive, destructive, or move into exploit-level testing are gated behind human authorization.

The goal is not to assume the AI will always make the right decision. The goal is to design the environment so that one bad decision does not automatically become one bad action.

We want our testing agents to be curious. We want them to notice strange things, propose unexpected tests, and investigate. But we want that creativity inside a very real box.

 

Logging Matters More Than I Expected

Another lesson from this work is how valuable complete observability becomes once an AI model is making decisions. When a traditional script fails, we inspect the code and the output. When an AI agent does something unexpected, the more interesting question is often: Why did it decide to do that?

Having the LLM request, the response, the selected command, the command output, the network activity, and the resulting evidence gives us an opportunity to reconstruct the entire decision path. In reviewing those paths, we have seen several different outcomes:

  • Sometimes the model was wrong.
  • Sometimes the tool output was misleading.
  • Sometimes the methodology needed improvement.
  • Sometimes the AI noticed something none of us expected it to notice.

All of those outcomes are useful during research. The point of logging is not simply compliance or accountability; it is also how we learn.

That has been a major part of our approach over the past nine months: let the agents work in controlled environments, observe them closely, understand where they succeed and fail, and improve the surrounding controls and methodology.

 

Responding When a New Critical Vulnerability Drops

One capability that has become especially interesting for Continuous Penetration Testing is rapid response to newly disclosed vulnerabilities. When a new critical CVE is announced, the first question is usually not does this vulnerability exist? The important question is: Which of our clients actually have the affected product exposed, and which of those systems appear vulnerable?

CASM-AI gives us a way to answer that much faster. Because CASM already maintains attack-surface and asset information, the platform can identify systems that appear to match the affected technology and dispatch targeted, non-destructive validation to the appropriate agents across the fleet.

We have been exercising exactly this kind of workflow. When a critical NetScaler vulnerability was disclosed, the platform was able to identify potentially affected client appliances and begin targeted version and exposure checks without waiting for a pentester to manually inventory every client first.

That is the kind of automation I find compelling. The AI is not deciding to exploit a new RCE across a client base. It is doing the repetitive work required to determine where a human pentester should look first. For a Continuous Penetration Testing program, shortening that gap from public disclosure to informed testing can be extremely valuable.

 

The Results Are Getting Hard to Ignore

The most important measure of any pentesting tool is whether it finds real things. At the time of this draft, SynerComm pentesters have human-validated more than 36 findings surfaced by CASM-AI, including 11 rated High Severity or greater.

That wording is intentional. The agents have surfaced substantially more potential issues than that, but I do not think a raw agent-generated finding count is the number we should brag about. A finding does not become meaningful just because an AI says it found one. What matters is that the platform is finding issues worth a pentester’s time and that experienced testers can reproduce and validate them.

That is the threshold I care about. If an AI system produces 1,000 “potential vulnerabilities” and a pentester has to spend two days proving that 995 of them are nonsense, we have not made the tester more efficient; we have created a new kind of scanner noise. The goal is the opposite: we want CASM-AI to do enough exploration, correlation, and preliminary validation that what reaches the pentester has a meaningful probability of being real.

The human tester still owns the final judgment:

  • Is it reproducible?
  • Does the evidence prove the condition?
  • Is it actually exploitable?
  • What is the realistic impact?
  • Can it be combined with something else?
  • How far should we go to demonstrate that impact safely?
  • What should the client do about it?

Those questions still require experience.

 

The Scanner Everyone Thought We Already Had

There is another part of this project that I find almost funny. For years, clients have sometimes assumed penetration testers had a magic scanner that basically performed the penetration test for us.

We never did.

Traditional vulnerability scanners are useful, and we still use them. They are good at finding known conditions they were designed to recognize. But they are not pentesters. They do not understand an application the way a person does, reason well across multiple observations, or notice something odd, get curious, and spend the next 20 minutes trying to understand why it happened.

That gap has always been enormous, and AI is starting to close it.

I would not describe CASM-AI as an automated replacement for a penetration test. We are not there, and I am not convinced that should even be the goal. But for the first time in my career, I can see a realistic path toward something much closer to the “scanner” people always assumed existed:

A system that can look at what it finds, think about it, and decide what to investigate next.

That is a very different capability.

 

Being Impressed Does Not Mean Being Careless

I am genuinely excited about what we are seeing, and I am also probably more cautious about it now than I was when we started. That is not a contradiction. The more capable these systems become, the more important it is to engineer the boundaries around them.

An AI that can reason through a web application, choose offensive tools, generate commands, and pursue an attack path is extremely useful. It is also something you should not deploy casually.

Our research has intentionally moved in stages: we test, review logs, validate results, adjust scope controls, break things in lab environments, watch how the agents behave, and make changes before giving them more freedom. That process is going to continue.

We do not need CASM-AI to be perfect before it becomes valuable. We need it to be useful, observable, constrained, and honest about uncertainty.

 

Where I Think This Goes

The long-term opportunity is not “AI replaces pentesters.” I think that is the wrong way to look at it.

The opportunity is that one experienced pentester, supported by well-designed AI agents, may eventually be able to investigate dramatically more of an environment than that same pentester could investigate manually. In practical terms, that can mean:

  • Broader coverage
  • More frequent testing
  • Faster follow-up on changes
  • Faster response when new critical vulnerabilities are disclosed
  • More time spent validating likely vulnerabilities
  • More time developing meaningful attack paths
  • Less time spent doing work that is necessary but repetitive

That is especially important in Continuous Penetration Testing, where the real challenge is not simply finding vulnerabilities once. It is keeping up with an attack surface that never stops changing.

We are still learning what CASM-AI can and cannot do well. But after nine months of working on this, I am no longer wondering whether AI will have a meaningful role in penetration testing. It already does.

The interesting questions now are how far we can safely take it, how much more capable it can make our pentesters, and how we make sure the quality of the human judgment at the end remains as strong as ever.

That is the part I am most excited to keep working on.