GLM-130B: An Open Bilingual Pre-trained Model

Benchmark Model Rank Results
language-modelling-on-big-bench-liteGLM-130B (3-shot)#1Accuracy: 15.11
language-modelling-on-big-bench-liteGLM-130B (1-shot)#2Accuracy: 14.91
language-modelling-on-big-bench-liteGLM-130B (0-shot)#3Accuracy: 13.31
language-modelling-on-lambadaGLM-130B (bidirectional attention)#7Accuracy: 80.2
language-modelling-on-the-pileGLM-130B#5Bits per byte: 0.634
language-modelling-on-the-pileJurassic-1#7Bits per byte: 0.65
language-modelling-on-the-pileGPT-3#16Bits per byte: 0.742
long-context-understanding-on-ada-levalChatGLM3-6b-32k#51k: 39.82k: 18.84k: 9.06k: 5.08k: 3.412k: 0.916k: 0.5
long-context-understanding-on-ada-levalChatGLM2-6b-32k#81k: 31.22k: 10.94k: 4.56k: 1.68k: 1.612k: 0.016k: 0.3
long-context-understanding-on-ada-leval-tsortChatGLM3-6b-32k#72k: 2.34k: 2.48k: 2.016k: 0.7
long-context-understanding-on-ada-leval-tsortChatGLM2-6b-32k#82k: 0.94k: 0.28k: 0.716k: 0.9
multi-task-language-understanding-on-mmluGLM-130B#23Average (%): 44.8