Autonomous penetration testing — SaaS platform

Autonomous penetration testing on your own infrastructure. Pick a target and launch the attack.

Authorized. Under your control. In your network.

attacker’s path
  • 01Reconnaissance
  • 02Foothold
  • 03Escalation
  • 04Lateral movement
  • 05Proof of impact
  • 06Report & remediation
WORKBENCH command center: waves, agents, findings, severity and run cost
Why now

Reactive defense can’t keep up.

The time from an exploit’s release to the first attack has shrunk from years to hours. The attack surface grows faster than a team can check it by hand — and traditional manual penetration testing is done once a year. Autonomous, agentic penetration testing (agentic AI) uses artificial intelligence to keep up: you run it as often as you need, and every time you get a real path to the objective — not a list of “maybes”.

Attack chain

A scanner sees individual bugs.
An attacker chains them into a path to the objective.

CRITICALMEDIUMLOW this is all a vulnerability scanner sees 010203040506

WORKBENCH doesn’t stop at a list of vulnerabilities. It chains them into a scenario, exploits them one by one and shows where they really lead — from the first foothold to critical business impact.

01

Reconnaissance

Maps subdomains, hosts, services and endpoints. The full picture before the first shot is fired.

02

Foothold

Exploits the first real vulnerability — evidence from action, not a version from a header.

03

Escalation

Elevates privileges and chains individual bugs into one working chain.

04

Lateral movement

Moves on to further systems, accounts and data — like a real attacker.

05

Proof of impact

Shows where the path really leads and what sits at the end of it.

06

Report & remediation

Every finding with evidence, a CVSS score and a concrete recommendation for the team.

From scope to report

Five steps. One run.

A single run goes from setting the boundaries all the way to a finished report. Below are screenshots from the live platform — the target is a deliberately vulnerable application in an authorized lab.

You start with boundaries

You set targets, exclusions and the exploitation level before the start. Denial-of-service attacks are disabled by default.

  • Targets and exclusions
  • Exploitation level
  • Privilege profile

Maps the attack surface

The swarm collects subdomains, hosts, services and endpoints, and the coverage matrix shows what has already been checked.

  • Asset inventory
  • Coverage matrix
  • Action queue
Attack surface: live services, coverage matrix and the queue of next actions

Proves, doesn’t guess

Every finding with a proof-of-concept, a CVSS score and a recommendation. A separate step screens out false positives.

  • A PoC for each one
  • CVSS and vector
  • “Verified” status

Shows the real path

Hosts, services, credentials and findings in one graph — exposure as paths, not table rows.

  • Attack graph
  • Risk rings
  • One view of the estate

Delivers a finished result

Executive summary, scope, methodology and findings with evidence — exported to a report ready for the client and the auditor.

  • Summary and scope
  • Findings with CVSS
  • Report export
Use cases

One engine. Many kinds of tests.

WORKBENCH runs tests autonomously, while you choose the scope and depth. From a single application to an entire infrastructure and cloud.

Application security

The full cycle: from a map of features to exploiting vulnerabilities in business logic.

Application securityPenetration testing →

Web applications and API

OWASP Top 10, access control, injection, IDOR and API vulnerabilities — including authenticated (greybox) testing with multiple roles.

Web · API · greyboxWeb application testing →

Network infrastructure

External and internal testing: from internet-facing exposure to lateral movement inside the local network.

Internal + externalInfrastructure testing →

Red team operations

Scenarios close to a real attack, carried out within agreed rules.

Red team opsRed Team →

Cloud testing

Configuration, identities and permissions in AWS, Azure and GCP — escalation paths in the cloud.

Cloud pentestCloud penetration testing →

Reconnaissance and OSINT

Mapping subdomains, assets and leaks — the full picture of the attack surface before someone else does it.

Recon & OSINTOSINT & reconnaissance →
Measured in the run
  • 0steps from scope to report
  • 0%critical success rate
  • 0%false positive check
  • ∞unlimited agent dispatch
Execution

An isolated execution runtime.

The whole platform runs on a dedicated server or virtual machine that you control. Tools are launched by an isolated Docker container (pentest-toolbox) or a runtime on the host — the backend only orchestrates them.

SHARED FILE SYSTEM — one store that every agent reads from and writes to. Findings, dumps and artifacts in one place.

“

Full autonomy and decision-making. The agent sets its own priorities based on signals from the engagement — what it has already found, what looks promising and where the path leads. When the first attempt fails, it adapts the plan and directs power to where it really leads to the objective. You set the boundaries — it picks the next move.

— the principle WORKBENCH is built on
Dedicated Workbench server / VM
BACKENDorchestrator — drives the run
controls
Toolbox runtime
Docker · isolated containerHost · runtime
nmap · nuclei · sqlmap · ffuf · httpx · katana
testing within ROE boundaries
TARGETYour network / application
Shared memory

Cybergraph.

Every host, service and finding from a run is stored and typed. The next run starts from what you already know — instead of scanning everything from scratch.

Traversal

Hostipv4 · dns
Contains
Serviceport · banner
Exposes
Endpointurl · method
Points to
FindingCVSS · evidence

Relationships map to MITRE ATT&CK techniques, and a finding points to a specific vulnerability. The graph links the whole estate into one searchable whole.

WORKBENCH — CYBERGRAPH · TABLE VIEW
TLS Configuration Assessmentfinding3
Login Form Security Analysisfinding4
104.21.5.243ipaddress—
Session Cookie Securityfinding4
https://app.example.test/registerurl—
Overly Permissive CORS Policyvulnerability4
CORS Reflection — curl testingevidence—
No Brute Force Protection — CRITICALfinding5
Interactive mode

Human in the loop. You decide.

Autonomy doesn’t mean “unsupervised”. In interactive mode the swarm asks about key decisions, and you can steer the run at any moment — with an ordinary conversation.

The agent asks before it acts

At key forks the swarm pauses and waits for your approval. You pick one of the suggestions or type your own command — nothing important happens without you.

  • Pauses at the intrusiveness threshold
  • Ready-made options or your own answer
  • Full decision context on screen

Steer the run with a conversation

The operator-agent chat works like a conversation with the person running the test. Ask about status and priorities, set the next actions, add or change the scope and authentication profiles — mid-run.

  • Status and priorities on demand
  • Change scope and auth profiles
  • Set the next actions on the fly
An offensive model — the leash in your hand

Autonomy that never slips out of control.

Offensive power only makes sense with hard guardrails. You set the scope, permissions and boundaries before the start — the platform enforces them throughout the run.

Human in the loop and over the loop

You see every step live. At any moment you take control without killing the run.

Group kill switches

Stop one operation or every operation in a given area with a single switch.

ROE boundaries

You set forbidden actions, the exploitation level and the privilege profile before the start. DoS disabled by default.

Sandbox on your host

Tools run in the pentest-toolbox container on a machine you designate — not on our backend.

Out-of-band testing

Optional validation over OOB channels (interactsh) confirms “blind” vulnerabilities.

ProxyGuard

The option to route agent traffic through a Burp or Caido proxy — full visibility and interception.

Persistent audit log

Every command, approval and artifact with a timestamp. You can replay the run step by step.

Your backend and database

You can keep the backend and data on your side — in a private cloud or on-premise.

Deployment

Deploy anywhere.

Nothing has to leave your network. The same engine runs as SaaS, in a private cloud and fully on-premise.

01

SaaS

backendwith us (VIPentest)
databasewith us
modelyour API provider
toolson your hosts

We host the platform. The tools still run on your side.

02

Private cloud

backendin your cloud
databasein your cloud
modelAPI or self-hosted
toolson your hosts

Credentials and findings stay in your environment.

03

On-premises

backendyour infrastructure
databaseyour infrastructure
modellocal (self-hosted)
toolson your hosts

Nothing leaves your network. A fully air-gapped option.

04

Our operator

operationour pentester
backendwith us
modelour choice
targetyours (authorized)

We run the test on your authorized target — you get the result and the report.

Plug in your stack

Your model. Your API cost. Your tools.

Any model backend

You connect your own account and pay for tokens directly on your side. You separately choose the model for the “mastermind” and for the specialists, as well as the reasoning effort level. We provide the platform, the agent swarm, reporting and updates.

  • Claude · Anthropic
  • OpenAI
  • OpenRouter
  • Local · Ollama / llama.cpp / vLLM

Built-in tools (pentest-toolbox)

Nmap
Nuclei
SQLMap
ffuf
httpx
Katana
Nikto
Amass
Subfinder
theHarvester
Trivy
Semgrep
Recon-ng
Shodan
Gobuster
WPScan
ldapsearch
NetExec (nxc)
Kerbrute
BloodHound
enum4linux-ng
rpcclient
smbclient
Impacket
Responder
Certipy
Evil-WinRM
Hashcat
Hydra
masscan

You tune the engine for the task: backend, mastermind model, specialist model, working mode, the number of parallel agents and the wave and time limits. The platform configures itself to keep costs in check.

Engine configuration: choosing the mastermind and specialist models, mode and parallelism
Three ways to start

Who WORKBENCH is for.

Security teams and CISOs

Run autonomous penetration testing on systems you designate yourself, in your own environment. Check your exposure more often than your team’s calendar allows.

Request a demo

Pentesters and red teams

Hand reconnaissance and repetitive exploitation attempts to the swarm, and focus on attack chains and interpretation. More engagements with the same team.

Try it

Software houses and IT providers

Build repeatable security testing into your release cycle and show clients proof that the application was tested offensively — on your infrastructure.

Let’s talk
FAQ

Frequently asked questions

Didn’t find your answer? Write to us — we respond within one business day.

Ask your own question
Can this tool disrupt production systems?

Yes, we cannot rule it out. WORKBENCH was built for maximum effectiveness and efficiency, so there is no way to guarantee that the actions taken by the agents will not disrupt a system. That is why on production systems we recommend working in interactive mode (human-in-the-loop) — the agent pauses and waits for your approval before critical steps — rather than continuous mode. In addition, the default profile is non-destructive (proof-only), and denial-of-service attacks remain disabled.

How is autonomous penetration testing different from a vulnerability scanner?

A scanner matches software versions against a database of known bugs and returns a list of “possible” issues, some of which do not exist. WORKBENCH runs an autonomous test: a swarm of AI agents maps the attack surface, forms hypotheses, tries to actually exploit vulnerabilities and chains bugs into scenarios. Every finding gets a proof-of-concept, a CVSS score and a recommendation. That is the difference between a list of warnings and a confirmed path to the objective.

Does it replace a human pentester?

No. The platform accelerates and scales your team’s work: it takes over reconnaissance, repetitive exploitation attempts and report assembly, while the operator controls the scope, approves actions and interprets the result. A human sees every step and can take control or stop the run at any moment.

Which AI model does the platform use and who pays for the API?

You choose the model. You get WORKBENCH as a SaaS license — you connect your own provider: Claude (Anthropic), OpenAI, OpenRouter or a local model (Ollama, llama.cpp, vLLM). You settle token costs directly with your provider. We deliver the platform and updates; you pay for the API only for what you actually use.

Can I run the platform in my own infrastructure or offline?

Yes. The execution tools run in a container (pentest-toolbox) on a host that you designate yourself — not on our backend. You can deploy the platform as SaaS, in a private cloud, fully on-premise, and with a local model even in an environment cut off from the internet.

How does the platform enforce the test boundaries (ROE)?

Before the start you define the Rules of Engagement: targets and exclusions, forbidden actions (denial-of-service disabled by default), exploitation level and privilege profile. Above the agreed level the agent pauses and waits for your approval. You also get group kill switches and a persistent audit log with every command.

Do findings include evidence and are they verified?

Yes. Every finding includes a proof-of-concept, the affected components, the outcome and a remediation recommendation, and a separate step screens out false positives. Unconfirmed vulnerabilities are flagged, and confirmed ones get a “verified” status.

Is it safe for production environments?

The platform is intended solely for testing targets that you authorize yourself, and it operates within ROE boundaries. The default profile is non-destructive (proof-only), DoS is disabled, and secrets touched by the agent are masked before they reach a log or report.

What does licensing and pricing look like?

We sell WORKBENCH on a SaaS model as a license for the platform; you cover the API cost of your own model separately with your chosen provider. We tailor the license scope to the number of operators and the deployment method. Write to us for a quote and demo access.

Time is running out. For attackers too.

We’ll show you the WORKBENCH platform on a sample, authorized target and choose the deployment model together — SaaS, private cloud or on-premise. You’ll also get a license quote and a list of what we need to get started.

  1. Request confirmationWe reply within 24 h on business days and ask about your target and environment.
  2. A call and a demo under NDAWe show the platform on an authorized target and agree on the deployment model.
  3. License quote and accessWe tailor the scope to the number of operators and the deployment method. No obligation.
  • Your AI model, your API cost
  • SaaS / on-premise deployment
We reply within 24 h

Write to us: platform demo and quote

Briefly describe what you want to test and how you want to deploy the platform. We respond within one business day.

    I consent to the processing of my personal data by VIPentest sp. z o.o. in order to respond to my enquiry. Details in the Privacy Policy.

    No commitment. NDA before any scoping call.

    Autonomous pentest · SaaS with your modelDemo