
AI models are often judged by benchmarks. That is useful, but it is not always enough.
A benchmark can show whether a model recognizes a known vulnerability pattern, explains a security concept, or classifies a code snippet correctly. What it cannot fully show is whether the model can help investigate unfamiliar software, follow a realistic attack path, and contribute to a finding that survives public disclosure.
Security research is a useful test case because the bar is higher than producing a plausible answer. A real vulnerability has to be reachable, reproducible, impactful, and fixable. It has to be explained clearly enough for maintainers to validate and patch it.
AI security researcher Sai Teja Erukude’s recent work offers one example of that kind of field validation. Before investigating several open-source Python projects, Erukude trained two specialized local LLMs for vulnerability research and CVE triage. Used in a human-supervised workflow, those models contributed to four disclosed and patched remote code execution vulnerabilities.
The findings were:
- CVE-2026-47117 – OpenMed – CVSS 9.8 Critical
- CVE-2026-47103 – Python StateMachine – CVSS 9.8 Critical
- CVE-2026-9147 – uproot – CVSS 7.8 High
- CVE-2026-10036 – SpeechBrain – CVSS 8.8 High
Each was manually reproduced, responsibly disclosed, assigned a public CVE, and fixed by the affected project.
Why benchmarks only tell part of the story
Benchmarks are valuable because they make comparison possible. They help answer whether one model performs better than another on a defined task. But vulnerability discovery is not a single defined task. It involves code reading, threat modeling, exploitability analysis, version review, and judgment about whether a behavior crosses the line from risky to reportable.
A model can identify a dangerous function and still be wrong about exploitability. A scanner can flag suspicious code that is unreachable. A model can also miss the importance of a harmless-looking metadata field that later becomes part of code generation, model loading, or deserialization.
That is why real-world validation is different. A disclosed CVE is not just an internal test result. It means the issue was concrete enough to be reviewed, fixed, and recorded publicly.
For applied AI, that matters. The question is not only whether a model can answer security questions. The stronger question is whether it can contribute to a workflow that produces externally validated outcomes.
How the models were used
The workflow started with static analysis. That stage surfaced risky primitives, suspicious data flows, and candidate execution paths.
The local LLMs were used after that initial screening. One model, vulnerability-researcher, analyzed and prioritized candidates. Its role was to reason about source code context and decide which signals deserved deeper human review.
The second model, cve-expert focused on exploitability, severity, affected components, and whether the issue appeared to meet the bar for CVE-level treatment. This created a useful separation between discovery and final triage.
That separation is important because security research has different stages. Finding something suspicious is not the same as proving impact. A useful workflow needs both curiosity and skepticism. The models helped form and challenge hypotheses. They did not make the final decision.
Human validation remained the gate. Every confirmed issue still required manual reproduction, safe testing, affected-version analysis, responsible disclosure, and maintainer coordination.
Four RCEs, different routes to execution
The four confirmed findings were all remote or arbitrary code execution issues, but they did not come from one repeated signature.
In OpenMed, the issue involved model loading behavior. A user-controlled model name could influence a path that loaded remote model code, creating a route to execution in affected versions.
In Python StateMachine, the issue involved unsafe evaluation of SCXML document expressions. A malicious document could reach Python evaluation behavior and execute code in the context of the hosting process.
In uproot, the vulnerable behavior came from dynamically generated Python reader classes. Metadata from a crafted ROOT file could be inserted into generated source code in a way that enabled code injection.
In SpeechBrain, the issue involved checkpoint metadata parsing. A malicious `CKPT.yaml` file could trigger unsafe YAML behavior during checkpoint discovery, even if the malicious checkpoint was not ultimately selected.
These are different technical paths. One involves trusted remote model code. One involves document expressions. One involves generated source. One involves unsafe metadata deserialization.
The shared pattern is broader: external data crossed a trust boundary and became executable behavior. That is a difficult pattern to evaluate with simple benchmarks. It requires reasoning about how software turns input into action.
What this shows about applied AI
The important result is not that AI found four bugs on its own. It did not. The important result is that specialized local LLMs helped focus a human-supervised vulnerability research workflow. They helped prioritize candidates, reason about exploitability, and support the path from static signals to confirmed impact.
That is a more realistic picture of applied AI in security. The models were useful because they worked inside a controlled process, not because they replaced expert judgment.
This distinction matters for maintainers too. Open-source projects are already under pressure from low-quality automated reports. If AI-assisted security research simply increases noise, it creates more burden than value.
A responsible workflow has to do the opposite. It should reduce false positives, preserve human accountability, and send maintainers reports that are specific, reproducible, and actionable.
A better standard for evaluating security AI
Security AI should be evaluated by more than model scores. The practical questions are harder.
Can the workflow find issues in real software? Can it distinguish suspicious code from exploitable code? Can the finding be reproduced? Can maintainers verify it? Does the process lead to a patch rather than just a warning?
Those questions are closer to the real value of applied AI.
Sai Teja Erukude’s findings show why field evidence should be part of the conversation. Public CVEs, maintainer fixes, and affected-version analysis provide a stronger validation signal than a private demo or benchmark result alone.
Benchmarks will remain useful. But in vulnerability research, the most meaningful evidence often comes from real code, real disclosure, and real remediation.
A patched CVE is not just a security outcome. It is also proof that an AI-assisted workflow produced something the software ecosystem could act on.
About Sai Teja Erukude
Sai Teja Erukude is a data scientist and AI security researcher with over seven years of experience building production AI, data, and software systems across industry and academia. Based in Wichita, Kansas, his work focuses on responsible and trustworthy AI, open-source vulnerability research, and applied data science. He has authored numerous research publications and serves as a peer reviewer for leading Q1 journals in AI, security, and applied data science. He holds an M.S. in Computer Science from Kansas State University.
Reference Links



