Between two pentests, hundreds of changes land in production. Most of them are never tested against the running app. Scanners run on every one of them and report possibilities, not proof. That's the gap D3 is meant to fill: short, safe, change-triggered checks between the human-led tests, with the annual pentest still in place.
Prove what's reachable.
Fix what matters first.
Scanners hand you hundreds of possible problems. D3 runs safe, authorized checks on your live application after each meaningful change, connects code, cloud and runtime context, and shows which weaknesses an attacker could actually reach. Then it confirms the fix worked.
A pile of findings. No proof.
- SAST
412 findings. 37 marked high.
- Dependencies
Critical CVE in a library you might not even call.
- Cloud config
38 misconfigurations. Which ones are exposed?
- Last pentest
Nine months ago. Forty deploys since.
- Slack, 4:52 pm
Is the invoices endpoint reachable without auth? Anyone sure?
Which ones can actually be reached.
Illustrative data. The product doesn't exist yet.
Pentests happen every 6 to 12 months. Code ships every day.
From a code change to a verified fix.
One scoped application. Checks that run when something meaningful changes. A record from the first signal to the confirmed fix.
Point D3 at one web application.
Give it the repository, the deployment and cloud configuration, and an authorized URL. Scope is written down before anything runs. Nothing outside it is touched.Scoped to one appRepository
source, routes, auth middlewareCloud config
IAM, network, storage policiesLive URL
staging or production, authorizedapp.example.com
1 application · authorized in writing · out of scope: everything else
Run after meaningful changes, not on a timer.
A merge, a deploy, a new route, a changed IAM policy. D3 looks at what changed and decides whether it touches anything security-relevant. Everything else is skipped.CI/CD hook- merge: add /api/invoices exportrun
- deploy: v2.14.0 to productionrun
- iam: widen s3:GetObject on reports bucketrun
- chore: bump eslint, fix typosskip
- ci: cache node_modulesskip
Plan authorized checks from context.
The AI reads the diff and the configuration, maps it to how the live app behaves, and picks the checks worth running. Every check carries the reason it was chosen.AI plans, you approve scopeinvoices.ts: new GET /api/invoices/{id}, tenant check removed in refactor
Route is public behind the load balancer. IDs are sequential. Invoices hold PII.
- Read one invoice from a second test tenant. Expect 403.
- Check the export endpoint for the same gap.
- Confirm the reports bucket policy didn't widen reach.
Run safe, non-destructive checks on the live app.
Checks are read-only or reversible. No data is destroyed, no accounts get locked, no load is generated. If a check can't be made safe, D3 reports it as an unconfirmed hypothesis instead of running it.Non-destructive- Read-only requests with test accounts
- Reversible writes in dedicated test tenants
- Header, auth and configuration probes
- Rate-limited, logged, attributable traffic
- Deleting or corrupting data
- Locking real user accounts
- Denial of service or load tests
- Anything outside the written scope
Show the attack path, with evidence.
A finding becomes a path: where it starts, what it reaches, and the request and response that prove it. Findings with no reachable path are ranked down, not deleted.Evidence-backed- 1Internet
no auth required to hit the route
- 2GET /api/invoices/{id}
object-level check missing
- 3invoices table
returns rows from other tenants
- 4Evidence
request, response, diff line, timestamp
409 other findings: no path found with current checks. Kept, ranked lower.
Give a specific fix. Confirm it worked.
Remediation points at the file, the config key or the policy. When the fix lands, the same checks run again and the path is marked closed or still open.Retest workflowinvoices.ts:41 · load invoice by (tenantId, id), not id alone
- retest · path 1 · after fix 9c21e0dclosed
- retest · path 2 · export endpointstill open
- retest · path 3 · reports bucketclosed
A lot of this market already exists. Here's the honest map.
These tools are good at what they promise. The open question I'm testing is whether a narrow, application-first validation loop is something teams still need between them.
| Category | Main alternatives | What they already promise | What D3 would add(featured) |
|---|---|---|---|
| Autonomous pentesting | Horizon3.ai (NodeZero), Pentera | Autonomous testing, attack-path validation, proof of exploitability, remediation and retesting. | A narrower, application-centric wedge: safer change-triggered validation that uses code, API, cloud and live-app context. |
| Continuous validation and BAS | Cymulate, Picus | Continuous control validation, breach-and-attack simulation, automated red teaming. | Focus on whether the deployed application itself is reachable and exploitable, rather than on validating security controls. |
| Dynamic AppSec (DAST) | StackHawk, Invicti, Burp Suite Enterprise | Dynamic web-app scanning, often wired into delivery workflows. | Move from “scan the application” to “prove the real attack path and validate the fix”. |
| Attack-path and exposure management | XM Cyber, Wiz | Model and prioritize exposure paths across cloud, identity, network and assets. | Execute carefully scoped validation instead of only modeling risk from configuration and telemetry. |
| Human-led testing | HackerOne, pentest firms, internal red teams | Human expertise, exploit validation, compliance reports, retesting. | Faster, recurring, lower-cost validation between human pentests. The humans stay. |
Safe first. Then useful.
A tool that pokes at live systems has to earn trust before it earns a budget line. These are the constraints the design starts from.
Authorized targets only.
D3 only touches an application its owner has put in scope, in writing. There's no mode for testing something you don't own.
Non-destructive by default.
Checks are read-only or reversible. If a check can't be made safe, it's reported as a hypothesis for a human, not run.
Evidence before claims.
A finding without a request, a response and a line in the code is a guess. Guesses get ranked down, not shipped as alerts.
Humans stay in the loop.
D3 is meant to sit between pentests, not replace them. Scope, risky checks and fixes are reviewed by people.
It plans and connects. It doesn't freelance.
The useful part is going past a one-line recommendation: planning authorized checks, linking evidence across code, cloud configuration and live behavior, and producing remediation that's specific enough to act on.- Plans which authorized checks are worth running after a change
- Connects evidence across code, cloud configuration and live behavior
- Writes remediation that names the file, key or policy to change
- Attack systems on its own or expand scope
- Summarize scanner alerts and call it analysis
- Decide what's risky enough to run without a human rule
Validating the problem before building the product.
I'm building, but customer evidence sets the direction. I have the technical background to build this. What I don't have yet is enough proof that teams want it. That comes first.
A scoped system for one web application. It uses source and deployment context to run safe checks after meaningful changes, identifies likely attack paths, and gives the team evidence plus a retest workflow.
- Thesis. Written down. Tested against the market map.
- Problem interviews, now. Talking to practitioners and professors. You're here.
- MVP. One web app. Safe checks after changes. Evidence and a retest workflow.
- Pilot. A few teams, real deploys, measured against their current process.
The questions that come up first.
Is this just another scanner?
Scanners list what might be wrong. D3 is meant to take that list, plus the code and cloud context, and test a live, authorized app to find out which items an attacker can actually reach. The output is a path with evidence, not a longer list.
Will it break my production app?
The design starts from non-destructive checks: read-only requests, reversible writes in test tenants, no load, no account lockouts. Anything that can't be made safe is reported as a hypothesis for a human instead of being run. You'd also be free to point it at staging first.
Does this replace our pentest or bug bounty?
No. Pentesters and bounty hunters find things a narrow automated loop won't. D3 is aimed at the months between those engagements, when code keeps shipping and nobody has re-tested the live app.
Do I have to integrate everything to start?
The MVP idea is one web application with three inputs: a repository, the deployment or cloud config, and an authorized URL. Less context means fewer checks and more unknowns, but it should still run.
Is it built yet?
No. I'm validating the problem first and letting conversations with security teams set the direction. The demos on this page are illustrations of the intended workflow, not screenshots.
Who's behind this?
One person with a Computer Engineering background, currently in an entrepreneurship program at Brown and a cybersecurity course at Harvard. The founder note below has the rest.
I don't want to build from assumptions.
I'm exploring a cybersecurity product for continuously validating the security of live applications. Most security tools scan code, dependencies or configurations before deployment and produce many possible findings. The question I'm investigating is whether teams need a better way to test an authorized live system, combine that with code and cloud context, and identify which weaknesses are actually reachable and worth fixing first.
I have a Computer Engineering background. I worked around fragmented security tooling in my Capstone, and I'm now using my entrepreneurship program at Brown and a cybersecurity elective at Harvard to test whether this is a real problem worth building around. I've started speaking with practitioners and professors because I'd rather be told I'm wrong now than find out after a year of building.
I'd start with safe, non-destructive testing and evidence-backed remediation. The goal is to help security teams find and validate real exposure earlier, then confirm the fix worked.
Enricco Gemha
Tell me how you validate security today.
I'm looking for 30-minute conversations with AppSec engineers, security leads and pentesters. No pitch, no deck. I want to understand what you do between pentests and what you'd want proven before you'd trust a tool near production.
- How do you decide what to fix first today?
- What happens, security-wise, after a deploy?
- When did you last confirm a fix actually closed the hole?
- What would you need to see before trusting a tool on a live app?