Cortico launches MedSafe-Dx to test AI clinical safety
Cortico has released MedSafe-Dx, a free open benchmark that measures whether frontier AI models are actually safe for clinical decision support, not just strong on medical exams. In the launch paper, the safest model still over-escalated most routine cases, while the weakest missed emergency cases, underscoring a gap between accuracy and safe judgment.
Why it matters: - AI models can score well on medical knowledge tests and still make unsafe clinical recommendations. - MedSafe-Dx is designed to measure the behaviors that matter in real care settings: escalation, false reassurance and uncertainty. - The benchmark gives hospitals, clinicians and buyers a way to compare safety, not just recall.
What happened: - Cortico launched MedSafe-Dx, a free, open benchmark for testing whether AI models are safe to support clinical decisions. - The launch paper evaluated 11 frontier large language models from OpenAI, Anthropic, Google and DeepSeek. - The models were tested across 250 simulated patient cases from the DDXPlus dataset. - The scoring was deterministic and did not use an LLM-as-judge. - The live leaderboard has since expanded to 12 models across six labs, including Meta and xAI.
The details: - MedSafe-Dx tests three safety behaviors: escalation sensitivity, avoidance of false reassurance and uncertainty calibration. - Escalation sensitivity checks whether a model escalates care when a condition could be fatal if missed. - Avoidance of false reassurance checks whether the model avoids an unsafe all-clear when a patient is at risk. - Uncertainty calibration checks whether the model reflects ambiguity appropriately when the clinical picture is unclear. - The safest model in the launch paper, GPT-5.2, reached a 97.6% safety pass rate. - GPT-5.2 still over-escalated 71% of routine cases. - Gemini 3 Pro Preview had the highest Top-3 diagnostic recall at 87.2%, but the lowest safety pass rate at 62.4%. - The worst-performing model missed 26 of 156 urgent cases. - Even GPT-5.2 missed five urgent cases. - The launch paper said the real risk is a wrong answer delivered in confident, medically fluent language. - A May 2026 report from Ontario's Auditor General found major flaws in 20 government-approved ambient AI scribes. - The report said 60% produced inaccurate clinical notes, 45% invented treatment plans or physical findings and 85% missed critical mental-health details. - The source also noted evidence that clinicians with AI-literacy training can reason less accurately after exposure to flawed AI output. - Clark Van Oyen, Cortico's CEO and co-founder, said high exam scores show textbook recall, not safe clinical advice. - Van Oyen said Cortico built MedSafe-Dx to give clinicians and health systems transparent, verifiable safety metrics. - Cortico said a pre-print is available on medRxiv. - Cortico also said the code, datasets and live leaderboard are available at msdx.cortico.health.
Between the lines: - The release argues that diagnostic accuracy and clinical safety are not the same thing. - The data suggest some models may look strong in benchmark settings while still failing at risk-sensitive judgment. - The reference to EHR alert fatigue points to a broader concern: excessive escalation can also reduce trust and usefulness. - The benchmark is meant to shift procurement decisions toward verifiable safety testing.
What's next: - Cortico is pushing AI buyers to ask vendors for peer-reviewed safety testing, not only exam-style accuracy results. - The open leaderboard gives the benchmark room to expand as more labs submit models. - The pre-print and shared materials may invite outside validation and follow-on research.
The bottom line: - MedSafe-Dx says the key question for clinical AI is not whether a model knows medicine, but whether it can practice it safely.
Disclaimer: This article was produced by AGP Wire with the assistance of artificial intelligence based on original source content and has been refined to improve clarity, structure, and readability. This content is provided on an “as is” basis. While care has been taken in its preparation, it may contain inaccuracies or omissions, and readers should consult the original source and independently verify key information where appropriate. This content is for informational purposes only and does not constitute legal, financial, investment, or other professional advice.
Sign up for:
World Publishing Review
The daily local news briefing you can trust. Every day. Subscribe now.
Check Your Email!
We sent a one-time activation link to: .
Confirm it's you by clicking the email link.
If the email is not in your inbox, check spam or try again.
Welcome back!
is already signed up. Check your inbox for updates.