A Systematic Study and Comprehensive Evaluation of ChatGPT on Benchmark Datasets

Benchmark Model Rank Results
natural-language-inference-on-anli-testChatGPT#6A1: 62.3A2: 52.6A3: 54.1