Specialized AI agent Mano-CUA outperforms Claude and Gemini on browser automation benchmark
A specialized vision-based AI agent called Mano-CUA 1.1 scored 41.7 on WebRetriever Protocol I, a benchmark testing real-world browser navigation and form-filling tasks. By comparison, Gemini 2.5 Pro Computer Use scored 40.9 and Claude 4.5 Computer Use scored 31.3, highlighting a notable performance gap. Unlike general-purpose models that rely on DOM parsing or accessibility APIs, Mano-P uses a pure vision-driven approach, interpreting screenshots the way a human would rather than reading page structure. This method proves more resilient on complex, dynamic web applications where page layouts shift after each user interaction. The model can run entirely on-device on Apple Silicon hardware with 32GB RAM, keeping user data local without sending screenshots to external servers.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in