My grader got every real answer right. Four red teams still broke it.
I built a grader that checks whether a quantum circuit an AI agent wrote is the circuit it was asked for. It uses no model in the loop. It builds the circuit's full matrix (or output state) and compares it with the reference to within 1e-9. Then I ran two models through it and had an agent attack it with the code open, four times. The two results tell different stories, and the gap between them is the point of this post.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in