Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
Open paper
Benchmark
Model
Rank
Results
arithmetic-reasoning-on-gsm8k
GaC(Qwen2-72B-Instruct + Llama-3-70B-Instruct)
#12
Accuracy: 90.91
question-answering-on-triviaqa
GaC(Qwen2-72B-Instruct + Llama-3-70B-Instruct)
#7
EM: 79.29
Rank counts only results with a code link.