Three Vision AI Models Tested on Product Catalogs: Relationship Mapping Remains Unsolved
A hands-on benchmark evaluated Mistral OCR, DeepSeek V4 Flash Vision Exp, and Qwen3-VL-32B-Instruct on a 138-page image-only product catalog to test their ability to convert visual data into structured, queryable records. While all three models accurately identified campaign prices, none successfully resolved shared price blocks or correctly extracted rotated SKU codes from a key test page. DeepSeek hallucinated incorrect product codes, while Mistral misclassified a list price as a campaign price, highlighting that character recognition is not the core challenge. A notable finding was that Qwen3-VL's extraction was actually successful, but a JSON field naming mismatch in the automated evaluator falsely reported a 0% accuracy score. The team concluded that evaluation pipeline integrity and human-in-the-loop workflows are as critical as model performance when building production-grade catalog extraction systems.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in