AI Chatbots Boost Rural Diagnostics but Mislead Lay Users, Studies Find

Two studies published in Nature Health on 6 February 2026 found that large language models significantly improved diagnostic accuracy among community health workers in Rwanda and physicians in Pakistan, with one trial showing doctor scores rising from 43% to 71% when using GPT-4o. Across all measured metrics, the AI models outperformed local clinicians in the Rwandan districts tested. However, a separate Oxford University study published days later in Nature Medicine painted a contrasting picture, finding that when nearly 1,300 ordinary users consulted chatbots for medical guidance, condition-identification accuracy fell to 34.5%, worse than the 47% achieved by people using conventional search engines. Correct triage decisions were made by only around 43% of lay users after consulting the AI, and the models gave inconsistent advice to different users describing identical symptoms. Together, the findings highlight a sharp divide between AI-assisted clinical settings and unsupervised public use, raising unresolved questions about safety and regulation.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in