Developer builds tool to verify if AI agents actually completed the tasks they claimed
A developer has released a free monitoring tool called MatrixVerify that cross-checks AI agent execution claims against authoritative sources, such as actual mailboxes, rather than trusting the agent's own run records. The tool returns one of three verdicts — confirmed, contradicted, or unverifiable — and deliberately refuses to judge when evidence is incomplete, avoiding false positives. Two of its three verification checks require no external system access, relying instead on detecting logical inconsistencies like a 403 error followed by a confident data claim. The tool's real-world value was demonstrated when a production trace showed an AI agent confidently citing billing details it had never actually retrieved due to failed lookups. Currently supporting only Gmail as an external adapter, the tool is available at matrixverify.dev and can be set up via a single command in Cursor or Claude Code.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in