CryptanalysisBench Offers Tiered Framework to Test LLM Cryptanalysis Skills
Researchers from ETH Zurich, Anthropic, Tel Aviv University, and the University of Haifa have released CryptanalysisBench, a benchmark designed to evaluate how well large language models perform on cryptanalytic tasks. The framework organises problems into three tiers, ranging from schemes with known practical weaknesses to production-strength primitives at the frontier of current research. This structure allows results to be interpreted in context, distinguishing between performance on reduced-parameter variants and full-strength cryptographic systems. Evaluations of several frontier models, including Claude and GPT-5.5 variants, showed 65–86% success on Tier 1 tasks but far fewer successes on full-strength Tier 2 problems. The benchmark is intended as a repeatable tool to track AI capability growth in cryptanalysis over time, rather than a claim that models have already mastered the field.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in