| code-generation-on-wikisql | GPT-4 | #4 | Execution Accuracy: 75.8 |
| code-generation-on-wikisql | ProTrix-Coder | #6 | Execution Accuracy: 72.3 |
| code-generation-on-wikisql | ProTrix | #8 | Execution Accuracy: 67.4 |
| code-generation-on-wikisql | GPT-3.5-turbo | #11 | Execution Accuracy: 55.0 |
| code-generation-on-wikisql | Llama-2-7B | #13 | Execution Accuracy: 17.4 |
| code-generation-on-wikisql | CodeLlama-7B | #14 | Execution Accuracy: 17.3 |
| fact-verification-on-feverous | ProTrix | #1 | Label Accuracy: 75.6 |
| fact-verification-on-feverous | ProTrix-Coder | #2 | Label Accuracy: 71.4 |
| fact-verification-on-feverous | GPT-4 | #3 | Label Accuracy: 71.0 |
| fact-verification-on-feverous | Plan-then-Reason (Ours) | #4 | Label Accuracy: 65.8 |
| fact-verification-on-feverous | GPT-3.5-turbo | #7 | Label Accuracy: 61.0 |
| fact-verification-on-feverous | Llama-2-7B | #8 | Label Accuracy: 47.1 |
| fact-verification-on-feverous | CodeLlama-7B | #9 | Label Accuracy: 43.0 |
| question-answering-on-hybridqa | GPT-4 | #2 | ANS-EM: 64.1 |
| question-answering-on-hybridqa | GPT-3.5-turbo | #4 | ANS-EM: 55.1 |
| question-answering-on-hybridqa | ProTrix-Coder | #5 | ANS-EM: 45.1 |
| question-answering-on-hybridqa | ProTrix | #7 | ANS-EM: 42.9 |
| question-answering-on-hybridqa | CodeLlama-7B | #8 | ANS-EM: 28.5 |
| question-answering-on-hybridqa | Llama-2-7B | #9 | ANS-EM: 27.6 |
| question-answering-on-tat-qa | GPT-4 | #8 | Exact Match: 80.8 |
| question-answering-on-tat-qa | GPT-3.5-turbo | #9 | Exact Match: 59.1 |
| question-answering-on-tat-qa | ProTrix-Coder | #10 | Exact Match: 52.2 |
| question-answering-on-tat-qa | ProTrix | #11 | Exact Match: 50.1 |
| question-answering-on-tat-qa | Llama-2-7B | #12 | Exact Match: 28.7 |
| question-answering-on-tat-qa | CodeLlama-7B | #13 | Exact Match: 28.4 |
| semantic-parsing-on-wikitablequestions | GPT-4 | #4 | Accuracy (Test): 72.9 |
| semantic-parsing-on-wikitablequestions | Plan-then-Reason (Ours) | #11 | Accuracy (Test): 65.2 |
| semantic-parsing-on-wikitablequestions | ProTrix-Coder | #17 | Accuracy (Test): 57.8 |
| semantic-parsing-on-wikitablequestions | ProTrix | #19 | Accuracy (Test): 56.2 |
| semantic-parsing-on-wikitablequestions | GPT-3.5-turbo | #21 | Accuracy (Test): 51.8 |
| semantic-parsing-on-wikitablequestions | Llama-2-7B | #25 | Accuracy (Test): 21.4 |
| semantic-parsing-on-wikitablequestions | CodeLlama-7B | #26 | Accuracy (Test): 13.1 |
| table-based-fact-verification-on-tabfact | Plan-then-Reason (Ours) | #8 | Test: 83.5 |
| table-based-fact-verification-on-tabfact | ProTrix | #12 | Test: 71.6 |
| table-based-fact-verification-on-tabfact | GPT-4 | #13 | Test: 71.5 |
| table-based-fact-verification-on-tabfact | ProTrix-Coder | #14 | Test: 70.6 |
| table-based-fact-verification-on-tabfact | GPT-3.5-turbo | #16 | Test: 68.8 |
| table-based-fact-verification-on-tabfact | CodeLlama-7B | #20 | Test: 49.5 |
| table-based-fact-verification-on-tabfact | Llama-2-7B | #21 | Test: 48.6 |
| table-claim-verification-on-scitab | ProTrix | #2 | Macro-F1 (3-class): 0.4500 |
| table-claim-verification-on-scitab | ProTrix-Coder | #3 | Macro-F1 (3-class): 0.4120 |