SShortSingh.
Back to feed

Researcher tests LLM judges' susceptibility to misleading answer formats.

0
·2 views

A developer created a benchmark called Judge Bait to test how LLMs evaluate answers. The test presented models with correct and incorrect answers, making the wrong ones harder to spot through tactics like verbosity or false authority tags. Leading models from OpenAI, Anthropic, and Google performed perfectly, even on difficult questions. Smaller models like GPT-5.4 nano struggled, especially on computationally intensive tasks where answers were nearly identical. The test revealed that some models could be misled by direct instructions to the evaluator embedded in the answer.

Read the full story at DEV Community

This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)

Log in to join the discussion and vote.

Log in

Related stories

0
ProgrammingDEV Community ·

Voice data industry shifts focus from bulk metrics to individual speaker rights and provenance

The speech data market is moving beyond describing datasets solely by volume metrics like hours or languages. New industry practices emphasize tracking individual contributors' consent, licensing terms, and economic claims. Platforms like SpeechData.ai, Intispeak, and Lokah now document per-speaker metadata and commercial-use consent chains. This shift aims to preserve ownership control and ensure proper compensation when data is used or modified.

0
ProgrammingDEV Community ·

U.S. seizes domains used by China-linked group to scan critical infrastructure

The FBI and Department of Justice have seized seven domains used by the China-linked Flax Typhoon group. The domains provided access to scanning and intrusion platforms operated by Beijing-based Integrity Technology Group. These tools were used to target a U.S. power company, airports in Japan and Poland, and Taiwanese energy firms and universities. The court-authorized seizure aims to disrupt the group's activities, which included operating an IoT botnet. A joint advisory from seven countries has been issued to document the group's methods and tools.

0
ProgrammingDEV Community ·

Tech professional advocates applying software skills to industries like healthcare

A software developer expresses frustration with the tech industry's narrow focus on AI tools for productivity gains. They observe that sectors like healthcare still rely on outdated paper systems and old technology. The author suggests software skills could solve meaningful problems in fields like healthcare or food service. They desire more real-world teamwork and connection than remote tech work provides.

0
ProgrammingDEV Community ·

Open-source app TouchGrass uses AI to generate personalized outdoor activity quests

A developer has created TouchGrass, an open-source web application that generates personalized outdoor activity suggestions. The tool asks users for their available time, preferred activity type, difficulty level, and whether they're alone or with friends before providing step-by-step quests. Built with HTML, CSS, JavaScript and Python, the application can use the Qwen2.5 AI model locally via Ollama or fallback to pre-built quests. The project aims to help people reduce screen time and make outdoor activities more accessible. The code is publicly available under the MIT License for others to use, modify, and contribute improvements.