TL;DR
‍
Every vulnerability gets a verdict either way: exploitable here or not exploitable here, with the evidence attached. Here is how we prove it, and how often the proof holds.

Exploitability is a state of your environment, so you can only prove it from the inside.

Across millions of vulnerability instances that Miggo observed in production, runtime analysis alone rules out 99.3% on average (96.5% to 99.52% across customers), each with its evidence attached. Live verification by our AI Attacker raises that to 99.9% (98.1% to 99.97%). For a team facing 10,000 findings, that is the difference between 70 to investigate and 10 to fix.

Figure 1: A backlog of 10,000 findings, reduced at each stage by the average cumulative rate across customers. Runtime analysis leaves 70 findings open. Live verification rules out 60 of them and proves the remaining 10 exploitable.

Today's tools can flag what might be exploitable, but they don’t prove what is not. Below, we explain why, how Miggo does it, and how we measure whether those proofs hold.

Exploitability Lives in Your Environment, Not the Advisory

A CVE is a fact about code. Exploitability is a fact about your environment. Whether an attacker can breach you through a given CVE depends on four questions: 

  1. Does the vulnerable function actually run? 
  2. Can an attacker reach it with input they control? 
  3. Which controls sit in the way? 
  4. Is all of that still true after this morning's deploy?

None of these answers appear in the advisory, and none are visible from outside the application. Any approach that works from the CVE outward, or from the internet inward, is guessing at the facts that matter. That is exactly where today's most common approaches fall short.

Why Today's Approaches Stop Short of Proof

Prioritization Produces a Ranking, Not a Verdict

The first commonly-used approach is prioritization: take the scanner list, add EPSS, KEV, asset criticality and a dependency-graph reachability pass, then sort. That produces a ranking, not a verdict. A team sorting 4,000 criticals into a smarter order is still patching 4,000 criticals. In CSA’s 2026 State of Application and AI Security only 9% of organizations remediate critical or high-severity vulnerabilities in production within 24 hours. Even among those that do, 77% were breached through a known vulnerability.

Outside-In Testing Cannot Explain a Miss

The second approach is outside-in validation: fire an exploit, or an AI agent that writes one, at the external surface. This approach tests rather than scores, but a miss is ambiguous. The attack may have been blocked, the code may be unreachable, or the agent may not have tried the right payload. Outside-in testing also cannot see internal services, queues or the second hop a request takes.

Neither approach can prove a negative. A ranking has no "not exploitable" bucket, only "lower."  A failed outside-in test proves only that the test failed. Proving a negative requires seeing the application from the inside, which is where Miggo starts.

How Miggo Proves It: Inside-Out, then Outside-In

One runtime model, five analyses, every exploit variant, one attacker. The first three work from the inside: they observe what the application actually does and rule out every path that cannot carry an exploit. The attacker works from the outside, aimed by what the inside already knows. Together they prove what is exploitable and how. Just as importantly, they prove what is not.

Inside-Out: Reverse engineer an attack from the vulnerability out

It starts with a runtime context graph (Miggo AppDNA), built from distributed traces, profiles and, where installed, Miggo's sensor. It records which functions execute, what calls them and what data they touch. Because the model is observed rather than declared, it doesn't inherit an SBOM's optimism.

The second input is live vulnerability and exploit intelligence (Miggo Pulse). Every new CVE, published exploit and attack primitive under discussion is picked up as it appears, then expanded into its known exploits, payload variations and mutations, and the function-level behavior each one produces when it lands. AppDNA says which execution flows reach which functions. Pulse says what an attack on them looks like, which functions must execute for it to succeed, and every form that attack has taken in the wild.

Then the inside-out process begins: five analyses run against that model:

  1. Reachability. Does the vulnerable function execute here, and is there a path from attacker-influenced input to it? A library installed but never called is not reachable. A function called only by a nightly job with trusted input is reachable, but not from the internet.
  2. Exposure. Which entry points sit on that path, and who can reach them: the public, authenticated users, partners or internal services only?
  3. Controls. What already sits between the attacker and the function: WAF rules, gateway policies, input validation, auth checks? A control counts only if it is shown to block the exploit's required input.
  4. Path enumeration. To say "not exploitable here", every caller and input to the vulnerable function must be enumerated, not just the one in the proof of concept. Miggo combines observed runtime calls with the static call graph. If a path appears in the code but has never been seen running, Miggo cannot confirm that it is safe. That finding stays open and is flagged as a coverage gap rather than marked not exploitable.
  5. Drift. A verdict (of exploitable or not exploitable) is true only for the code running today. However, deploys add routes and rewire services every day. Miggo re-checks every verdict whenever the runtime model changes. If a verdict flips, for example from not exploitable to exploitable, Miggo shows what changed to cause it.

Outside-In: Live Verification by an Attacker That Knows Where to Look

The five analyses are passive. They observe the system from the inside, produce a hypothesis and utilize observation to reverse engineer and prove the attack path. Outside-in live verification is active. The AI Attacker probes from the outside, in an environment the customer approves (staging, a mirror or production), to turn that hypothesis into proof.

For positives, it reproduces the exploit along the open path in a controlled run. For negatives, it attempts the exploit and its variants against every enumerated path.

This is not outside-in testing in reverse. An outside-in agent that misses doesn't know whether the function is even there. The AI Attacker works from the runtime model, so it knows exactly where the function is and every path that reaches it. A failed attempt on a known, enumerated path is evidence, not silence.

A negative verdict requires both: every path closed by analysis, and every path resisting the attack.

This is where the last step comes in, going from 99.3% to 99.9%. Analyses leave about 0.7% open: paths they can't close on their own. The attacker rules out six in seven of those, and proves which ones are real.

Figure 2. How Miggo reaches a verdict. The five analyses rule out about 99.3% of findings, each with evidence attached. The AI Attacker tests the roughly 0.7% left open, rules out six in seven of those and reproduces the exploit on the rest, which are then mitigated and re-tested by the same attacker.

The Results: 99.9% Ruled Out, None Contradicted on Replay

Reduction is the share of observed vulnerability instances proven not exploitable where they run, averaged per customer. Findings left with a coverage flag count as not reduced. Negatives are checked against a full exploit-and-variant replay in a mirror of the environment, because a pentest that missed something proves nothing.

Measure Runtime analysis With live verification
Reduction, average 99.3% 99.9%
Reduction, range across customers 96.5% to 99.52% 98.1% to 99.97%
Negatives contradicted on replay 0% 0%
Measure:
Reduction, average
Runtime analysis:
99.3%
With live verification:
99.9%
Measure:
Reduction, range across customers
Runtime analysis:
96.5% to 99.52%
With live verification:
98.1% to 99.97%
Measure:
Negatives contradicted on replay
Runtime analysis:
0%
With live verification:
0%

Sample size: runtime analysis figures cover 6.3 million vulnerability instances observed in production over the past 8 months.

React2Shell in Practice: Nine Findings, Two Exploitable, Seven Proven Closed.

When React2Shell (CVE-2025-55182) was disclosed, one customer had nine findings on it. Reachability and exposure analysis put two in scope. Path enumeration closed the other seven, and the AI Attacker failed against every path into them.

On the two open instances, the AI Attacker reproduced the exploit. A mitigation was deployed and verified against the exploit and 12 bypass variants within 55 minutes. The patch landed five days later. Seven findings never became an emergency, and nobody had to take that on faith.

Why the Negative is the Half that Pays

In most environments, most critical findings on a given CVE are not exploitable where they're deployed. Each one proven closed is an emergency change that doesn't happen and an engineer who stays on the roadmap.

The few that are exploitable get the opposite treatment. The same attacker that proved the exploit then verifies the mitigation holds, so protection is in place in minutes while the permanent fix moves through the normal cycle. That is how the exploitability window shrinks toward zero from both sides: fewer emergencies, and no open exposure while teams argue over which findings are real.

A ranking tells a team where to look first. Proof tells them what can safely wait, and that is where the time comes back.

Put Miggo’s Technology to the Test

A benchmark is only as good as its ground truth, so we are publishing the definition and we invite the test. Bring one critical CVE with a public exploit and your current list of affected services. In a 30-minute session, we will run the five analyses against your environment, issue a verdict for each instance, and show the evidence behind every one, including the ones we mark as not exploitable. 

If you have a recent pentest or bug bounty report, bring that as well, and we will put our verdicts next to its findings.

Book your 30-minute session now

<script src="https://cdn.jsdelivr.net/npm/gsap@3.12.5/dist/gsap.min.js"></script>
<script src="https://cdn.jsdelivr.net/npm/gsap@3.12.5/dist/Flip.min.js"></script>

<script>
  document.addEventListener("DOMContentLoaded", (event) => {
    gsap.registerPlugin(Flip);
    const state = Flip.getState("");
    const element = document.querySelector("");
    element.classList.toggle("");
    Flip.from(state, {
      duration: 0,
      ease: "none",
      absolute: true,
    });
  });
</script>
<script src="https://cdn.jsdelivr.net/npm/gsap@3.12.5/dist/gsap.min.js"></script>
<script src="https://cdn.jsdelivr.net/npm/gsap@3.12.5/dist/Flip.min.js"></script>

<script>
  document.addEventListener("DOMContentLoaded", (event) => {
    gsap.registerPlugin(Flip);
    const state = Flip.getState("");
    const element = document.querySelector("");
    element.classList.toggle("");
    Flip.from(state, {
      duration: 0,
      ease: "none",
      absolute: true,
    });
  });
</script>