344
Productivity & Workflow355
Automation & Workflow225
Software Development251
Marketing & Growth192
AI Infrastructure & MLOps175
Writing & Content Creation203
Data & Analytics142
Photography & Imaging156
Design & Creative170
Customer Support133
Sales & Outreach125
Voice & Speech135
Education & Learning131
Operations & Admin87
AI is getting very high scores on medical knowledge tests, but researchers say these tests may no longer show how AI will perform in real clinics.
In short: AI is getting close to perfect scores on medical knowledge tests, but researchers and clinicians say real medical care still needs human judgment and oversight.
Some of today’s top AI systems can score extremely well on clinical knowledge benchmarks, which are standardized tests used to compare systems (like using the same exam to compare students). Stanford’s AI research group reported that OpenAI’s o1 model scored 96.0% on MedQA, a well known medical question set. Stanford also said MedQA may be nearing “saturation,” meaning the test is becoming too easy for the best systems and stops showing meaningful differences.
Researchers have also pointed to studies where AI beat doctors on certain narrow tasks. Examples include diagnosing difficult written case descriptions, spotting some cancers, and predicting which patients have a high risk of dying. But other reviews find the best results often come from teamwork, where a doctor and AI work together, instead of AI working alone.
At the same time, AI is spreading in day to day healthcare work. One big use is automatic medical note writing, where software turns a doctor’s conversation with a patient into a visit summary, like a fast assistant taking notes. Stanford reported broad adoption of these tools in 2025, and The New York Times described medical “scribes” as a strong foothold for AI.
Regulators are allowing more AI tools into healthcare, including 258 AI medical devices authorized by the FDA in 2025, many without new clinical trials. Watch for a growing gap between high test scores and safe real world performance, especially as consumer health chatbots make medical claims without authorization.
Source: NYTimes