Study tests LLMs' ability to spot UI rule violations from screenshots vs DOM data

A developer ran a controlled experiment to measure how accurately large language models can detect UI threshold violations — such as font sizes and tap target dimensions — using only screenshot crops, with ground truth labels derived from DOM measurements. Ten custom style rules were tested across 12 pages rendered at two viewport widths, producing 96 scorable judgment pairs after excluding pages where elements were absent. GPT (via ChatGPT UI) achieved 90% accuracy on the contrastive rules containing both pass and fail cases, while Gemini 2.5 Flash scored 68% on the same set. Both models performed well on visually apparent rules like wrapping headings, but struggled significantly on purely numeric rules such as font size and line length. The author notes the tap-target rules could not test true threshold discrimination, as each rule contained only violations or only passes across all tested pages.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in