ChatGPT and rivals put to the test: can AI match doctors in pre-surgery risk scoring?

NCT ID NCT07696221

First seen Jul 10, 2026 · Last updated Jul 10, 2026

Summary

This study tests whether four popular AI chatbots — ChatGPT, DeepSeek, Gemini, and Claude — can correctly classify surgical risk using anonymous patient vignettes. Researchers compare the AI's risk scores to those given by a panel of senior anesthesiologists. If successful, AI could help standardize preoperative assessments and reduce human variability.

What this could mean

Our plain-language read of the trial. This is informational only — not medical advice or a prediction.

Active substance
Large language models (ChatGPT, DeepSeek, Gemini, Claude)
What this could lead to
If these AI models prove reliable, they could help standardize preoperative risk assessment and reduce variability among clinicians.
What could go wrong
This is a small, retrospective study using anonymized vignettes, not real-time patient care. AI may not perform as well in complex or emergency cases.

This is an AI summary of the original study and may miss details. Read our disclaimer.

Get updates

Get notified about this study

Sign up to get updates when this study changes or when new studies for ANESTHESIA are added.

Our safety recommendation!

By submitting, you agree to our Terms of use

Conditions

The condition(s) this trial relates to.

As listed by the trial registrant

The condition terms exactly as the trial's registrant entered them.

Contacts and locations

Study contacts

  • Contact

    Phone: •••-•••-•••• Email: •••••@•••••

More trials for these conditions

Other studies related to the condition(s) this trial covers.