Training language models to follow instructions with human feedback

Benchmark Model Rank Results
question-answering-on-timequestionsInstructGPT#8P@1: 22.4
question-answering-on-tiqInstructGpt#5P@1: 23.6