Playbook

Evaluating Automated Vulnerability Detection - Research Report

A research report conducted in collaboration between Cytix and the University of Lancashire on vulnerability detection gaps across methodologies and CWE vulnerability classes.

5 min

Jacob Hothersall

Playbook

Evaluating Automated Vulnerability Detection - Research Report

A research report conducted in collaboration between Cytix and the University of Lancashire on vulnerability detection gaps across methodologies and CWE vulnerability classes.

5 min

Jacob Hothersall

Playbook

Evaluating Automated Vulnerability Detection - Research Report

A research report conducted in collaboration between Cytix and the University of Lancashire on vulnerability detection gaps across methodologies and CWE vulnerability classes.

5 min

Jacob Hothersall

In this article

No headings found on page
No headings found on page

Join our newsletter

Receive the latest advancements, playbooks, and industry insights in software change security understanding.

We'll store and process this information in line with our privacy policy to provide you our products and services. You may opt out of this at any time.

Abstract

Automated vulnerability scanners are widely used in security for modern DevSecOps pipelines to identify flaws and weaknesses in web applications. Despite their widespread adoption, their true detection capabilities and effectiveness across different vulnerability classes remains insufficiently understood.

Existing evaluations of automated vulnerability scanners are scarce, additionally being limited in both scale and detail. As a result, organisations rely on these tools without clear insight into their limitations.

Therefore, this research aims to close that gap by systematically analysing detection gaps and their sources across methodologies (such as SAST and DAST scanning) and CWE vulnerability classes.

These new contributions to research can assist working towards providing accurate and meaningful statistics across detection of the most common vulnerability types across a variety of testing frameworks, providing specific insights into scanner limitations and strengths; the sources of these detection gaps and how these vary across methodologies

Summary

*Recall = detected opportunities ÷ total opportunities.

13,112

Detection opportunities tested

27.7%

Overall recall (CWE-based)

81.9%

Best coverage - CodeQL (SAST)

0.5%

Worst coverage - Nuclei (DAST)

4.2x

SAST vulnerablity detection vs DAST

79.9%

DAST - DAST miss overlap

Key takeaways

No 'catch all' method

No single scanner, and no single methodology, provides anywhere close to full coverage. Even the strongest tool tested (CodeQL) still missed roughly 18% of known-vulnerable test cases; the weakest (Nuclei) missed over 99%.

SAST outperforms DAST

Static analysis (SAST) substantially outperformed dynamic scanning (DAST)- 45.5% vs 10.9% average recall due to scanning source code directly, rather than being limited to what is reachable and observable over HTTP.

More scanners doesn't increase coverage

Running more scanners doesn't close the gap as much as it might seem to: scanners rarely agree on what they detect, but overlap heavily in what they miss (up to ~80% miss-overlap between DAST tools). Coverage gaps are systemic, not random.

Cryptographic issues need more than automated scanners

The weakest spots consistently found across tools were cryptographic issues (weak hashing, insufficient randomness, risky algorithms), broken access control / trust-boundary issues, and several injection classes (LDAP, XPath, OS command). These deserve targeted manual review or supplementary tooling.

Humans and machines are still the best combination

Automated triage of ambiguous scanner findings (NLP-based pre-filtering) cut a human review workload from 108 unique candidates down to a manageable set, but only 13 were ultimately accepted as genuine detections after human review. Automation reduced effort but did not replace the need for expert sign-off.

The results

Scanner Coverage

Recall = detected opportunities ÷ total opportunities.

CodeQL and Semgrep (both SAST) accounted for the large majority of successful detections; all three DAST tools trailed well behind.

Where scanners consistently fall short

Ten weakness classes account for roughly 78% of all recorded misses across every scanner and benchmark tested.

Adding more scanners doesn't close the gap

Scanners rarely detect the same vulnerability, but they very often miss the same one, especially within DAST tooling. This means stacking multiple DAST scanners buys far less additional coverage than it might appear to.

Highest-priority blind spots (lowest recall, by weakness class)

These specific classes are the best candidates for a manual review step or a supplementary specialised tool, rather than relying on general-purpose scanner defaults.

Evidence-based review: automation helps triage, not decide

Where CWE ground-truth was missing or ambiguous (e.g. the WAVSEP and XBOW benchmark suites), findings were instead validated by hand against actual scanner evidence - payloads, responses, and error output.

NLP techniques were used only to pre-filter and prioritise. It reduced 183 raw candidate findings to 108 unique ones, but of those, only 13 survived full human review as genuine, benchmark-aligned detections (12 of the 13 fell within Injection or Broken Access Control).

The takeaway for us: automated triage is a useful time-saver, but the final call still needs a person who understands the application.

Benchmark composition

Three benchmark suites, OWASP Java, OWASP Python, and ASDF, have CWE ground truth and support a full recall breakdown below.

XBOW and WAVSEP don't, and are assessed separately through evidence based review; the evidence-based hit counts mentioned below all come from that same review, not from recall.

OWASP Benchmark (Java)

8,490 opportunities, 64.7% of the recall dataset. Recall by scanner:

CodeQL detected every opportunity in this suite. Also contributed 2 of the 13 evidence-based hits.

OWASP Benchmark (Python)

2,712 opportunities, 20.7% of the recall dataset. Recall by scanner:

CodeQL's recall drops from 100% on Java to 25.2% here, and ZAP overtakes it as the strongest scanner on this suite. No evidence-based hits came from this suite.

ASDF (Cytix)

1,910 opportunities, 14.6% of the recall dataset, run by five of the six scanners. CodeQL was excluded, due to unsupported languages in the test set, including PHP. Recall by scanner:

Combined recall across the five scanners tested here is the lowest of the three suites, around 7%, though it isn't the weakest individually for every scanner- Nuclei, for instance, actually performs best on ASDF. ASDF also produced 7 of the 13 evidence-based hits, the most of any suite.

The ASDF is an open source framework managed by Cytix for understanding the capabilities of automated detection methods at identifying classes of application security vulnerabilities. ASDF is designed to evaluate and compare the effectiveness of various security scanners in detecting common web application vulnerabilities. It provides a standardized set of vulnerable applications and a framework for testing security tools against them.

XBOW

Not part of the recall/opportunity dataset, since it lacks CWE ground truth. Contributed 3 of the 13 evidence based hits.

WAVSEP

Also assessed only through evidence-based review, contributing 1 of the 13 evidence-based hits.

Conclusion

This research investigated the effectiveness of automated vulnerability scanners via largescale benchmark-driven evaluation. By undergoing evaluation of six widely utilised scanners in industry across several benchmark suites, it became possible to analyse differences in detection capability, vulnerability coverage, scanner overlap and blind spots across both SAST and DAST-based methodologies.

The results demonstrated that the variety of scanner performance varies considerably on both the methodology and weakness class. One notable finding was that SAST-based scanners substantially outperformed DAST-based scanners in overall recall, largely due to their visibility of source code and internal application logic.

However, no scanner achieved complete coverage, with significant detection gaps remaining across all evaluated tools. On the other hand, the evidence-based validation workflow further demonstrated that vulnerability classification can be assisted via automated processing, but should not be relied on without any oversight.

By combining NLP-assisted filtering with human review, the workflow reduced 183 candidate findings to just 13 validated detections. Therefore, whilst automated processing proved useful for the organisation and prioritisation of scanner findings, human review with real-world contextual knowledge remains necessary to verify this evidence and keep classification accurate.

Overall, the research concludes that automated vulnerability scanners provide valuable assistance with workflows of modern cyber-security, but should not be treated as comprehensive security solutions.

Significant coverage gaps remain across all scanners, making relying on a single scanner or even numerous unlikely to provide sufficient accuracy. Further improvements in automated security testing will therefore depend not only on improving these individual scanners, but also on addressing frequent blind spots that continue to be shared across multiple tools, whether SAST or DAST.

Focsuing scanners attention

No scanner in this study caught everything. Adding more tools across every vulnerability just hands reviewers a longer list. So the useful question comes earlier: which software change deserve a scan at all?

Cytix reads each ticket, PR and release for context (the asset, the environment, the history) and qualifies whether it carries real risk against your own policy.

Your scanners and reviewers are then pointed at the changes that do, and every decision stays attached to the change that caused it. Security time goes where it changes the outcome.

See how Cytix focuses your scanning on the changes that matter.

Join our newsletter

Receive the latest advancements, playbooks, and industry insights in software change security understanding.

Join our newsletter

Receive the latest advancements, playbooks, and industry insights in software change security understanding.

We'll store and process this information in line with our privacy policy to provide you our products and services. You may opt out of this at any time.

Legal

101 Princess Street, Manchester, United Kingdom M1 6DD

© 2026 Cytix Ltd. All rights reserved.

Join our newsletter

Receive the latest advancements, playbooks, and industry insights in software change security understanding.

We'll store and process this information in line with our privacy policy to provide you our products and services. You may opt out of this at any time.

Legal

101 Princess Street, Manchester, United Kingdom M1 6DD

© 2026 Cytix Ltd. All rights reserved.

Join our newsletter

Receive the latest advancements, playbooks, and industry insights in software change security understanding.

We'll store and process this information in line with our privacy policy to provide you our products and services. You may opt out of this at any time.

Legal

101 Princess Street, Manchester, United Kingdom M1 6DD

© 2026 Cytix Ltd. All rights reserved.