Back to knowledge graph
science
machine-learning
Confidence 75%

Large language models encode clinical knowledge Abstract Large language models (LLMs) have demonstrated impressive capabilities, but the bar for clinical applications is high. Attempts to assess the clinical knowledge of models typically rely on automated evaluations based on limited benchmarks. Here, to address these limitations, we present MultiMedQA, a benchmark combining six existing medical question answering datasets spanning professional medicine, research and consumer queries and a new ...

Source:
Cited 3626 times
undefined | Awareness Public Knowledge