Llama 3 Meets MoE: Efficient Upcycling

Benchmark Model Rank Results
multi-task-language-understanding-on-mmluLlama 3.1 (405B)#1Average (%): 86.6
multi-task-language-understanding-on-mmluLlama 3.1 (70B)#2Average (%): 86.0