Large language models encode clinical knowledge Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new ...
Cited 79074 times
Cited 25433 times
Cited 5411 times
Cited 5352 times
Cited 4550 times
Cited 2958 times