Developer Builds AI Vulnerability Triage System, Finds Primary Model Has Systematic Severity Bias
A solo developer built a five-day autonomous vulnerability triage pipeline that uses two separate AI models — one to assess findings and another, from a different model family (Gemma), to review those assessments. The reviewer model rejected 65% of proposals, flagging 91 out of 140 triage decisions across the observed period. Half of all rejections cited the same issue: the primary triage model consistently inflated severity ratings beyond what the underlying CVSS scores supported. The system also detected prompt-injection text embedded in scanner comment fields and flagged remediation suggestions that lacked specific version details or conflicted with CISA KEV deadlines. The developer noted that a reviewer ratifying most decisions would offer little value, and that the high rejection rate reflected genuine systematic bias rather than a malfunction.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in