Evaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark

Benchmark Model Rank Results
named-entity-recognition-ner-on-ncbi-diseaseGPT-4F1: 65.98
question-answering-on-bioasqGPT-4Accuracy: 85.71
question-answering-on-blurbGPT-4Accuracy: 80.56
question-answering-on-pubmedqaFlan-T5-XXLAccuracy: 76.80