Study Finds AI Safety Training Flips Expected Behavior in Civil Unrest Simulations
Researchers tested a large language model's ability to simulate citizen decisions in a civil violence model. Using the Qwen 27B model, they found the AI's probability of choosing protest action decreased as social tension increased, opposite to the original model's prediction. When the choice labels were changed from 'act/wait' to 'protest/comply', the AI's response curve flipped to show increasing protest probability with tension. The researchers suggest AI safety alignment may be causing this anti-action bias and recommend calibration stages for such simulations.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in