Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models

Benchmark Model Rank Results
question-answering-on-feverCoA#1EM: 68.9
question-answering-on-feverSelf-Ask#3EM: 64.2
question-answering-on-feverDSP#5EM: 62.2
question-answering-on-feverCoA w/o actions#6EM: 54.2
question-answering-on-feverZero-shot#8EM: 50
question-answering-on-strategyqaCoA#3EM: 79.2
question-answering-on-strategyqaSearchChain#5EM: 77
question-answering-on-strategyqaCoA w/o actions#6EM: 70.6
question-answering-on-strategyqaLeast-to-Most#8EM: 65.8
question-answering-on-truthfulqaCoA#26EM: 67.3
question-answering-on-truthfulqaCoA w/o actions#28EM: 63.3
question-answering-on-webquestionsCoA#1EM: 70.7
question-answering-on-webquestionsCoA w/o actions#2EM: 64.7
question-answering-on-webquestionsDSP#4EM: 59.4
question-answering-on-webquestionsFew-shot#7EM: 44.7
question-answering-on-webquestionsZero-shot#10EM: 43
question-answering-on-webquestionsCoT#13EM: 42.5
question-answering-on-webquestionsReact#18EM: 38.3
question-answering-on-webquestionsSelf-Ask#21EM: 31.1
question-answering-on-webquestionsToT#25EM: 26.3