Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling

Benchmark Model Rank Results
arithmetic-reasoning-on-gsm8kGaC(Qwen2-72B-Instruct + Llama-3-70B-Instruct)#12Accuracy: 90.91
question-answering-on-triviaqaGaC(Qwen2-72B-Instruct + Llama-3-70B-Instruct)#7EM: 79.29