DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

Zhuoran Yu1, Le Thien Phuc Nguyen1, Jaden Park1, Xinyi Gu2, Zexue He3, Soochahn Lee4, Rogerio Feris5, Yong Jae Lee1

1University of Wisconsin-Madison   2Massachusetts Institute of Technology   3Stanford University   4Kookmin University   5MIT-IBM Watson AI Lab

Correspondence: zhuoran.yu@wisc.edu

Overview

DocHop is designed to evaluate whether multimodal models can perform integrated chart–document reasoning rather than treating charts and text as separate sources of information. Each instance is built around a symbolic reasoning trace that specifies how candidate entities should be filtered and combined through multi-step constraints. The document narrative verbalizes this reasoning specification, while the charts provide the corresponding numerical evidence. Questions refer to a semantic reference label introduced in the narrative instead of directly naming the target entities, forcing models to first resolve the relevant entities from context and then retrieve or aggregate evidence from the charts. In this way, DocHop isolates a controlled out-of-domain reasoning challenge: using document context to determine which chart evidence is relevant and how it should be reasoned over.

Dataset Preview

Leaderboard

News [2026-09] We added results for the latest models — GPT-5.6-terra, GPT-5.6-luna, Claude-Sonnet-5, and Gemini-3.1-Pro. Due to compute budget constraints we can't chase every new model release, and we welcome community-submitted results for models we haven't been able to cover.
Model Value Retrieval Counting Numeric Reasoning Ranking Hypothetical Fact Checking Overall
Human Eval
95.3694.1595.5393.5686.5496.6793.29
GPT-5.6-terraReasoningProprietary
95.7092.6986.5988.4288.8696.3691.22
GPT-5.6-lunaReasoningProprietary
89.4090.3583.8087.4687.7091.8288.33
Claude-Sonnet-5ReasoningProprietary
74.8374.5657.5469.4564.0483.3370.11
GPT-5.2-ReasoningReasoningProprietary
66.5667.8446.9360.7758.7078.7962.83
Gemini-3.1-Pro-ReasoningReasoningProprietary
53.3158.1928.7748.2338.2869.3948.55
Gemini-2.5-Pro-ReasoningReasoningProprietary
41.3949.1217.6039.5535.7363.3340.60
GPT-5.2Proprietary
40.7356.1419.8337.3035.5055.1540.36
Gemini-2.5-Flash-ReasoningReasoningProprietary
35.1040.9411.1728.6222.2758.4832.02
GPT-5-mini-ReasoningReasoningProprietary
29.1423.689.5031.1923.6755.7628.25
Gemini-2.5-FlashProprietary
27.1525.738.9428.3016.9446.3624.88
Qwen-2.5-VL-72BOpen-Source
23.8426.327.8225.0817.8748.7924.40
Qwen3-VL-8BOpen-Source
16.8922.5110.6132.1516.4749.7024.16
Qwen-2.5-VL-32BOpen-Source
21.8519.016.7022.5116.4747.2721.79
Qwen-2.5-VL-7BOpen-Source
16.8913.744.1924.1212.0646.9719.05
Claude-4.5-SonnetProprietary
4.6428.652.517.7210.4455.1517.94
Molmo-7B-O-0924Open-Source
4.3023.982.5112.2211.6046.9716.73
Claude-4.5-HaikuProprietary
2.9819.011.964.8211.3758.7916.35
Molmo-7B-D-0924Open-Source
9.9321.052.2313.1810.2141.5216.01
InternVL-3.5-38BOpen-Source
4.9721.934.193.869.0550.6115.57
InternVL-3.5-30B-A3BOpen-Source
2.3219.302.792.8911.8349.3914.75
InternVL-3.5-8BOpen-Source
3.3110.231.964.828.5853.0313.45
LLaVA-Next-LLaMA3-8BOpen-Source
0.009.650.560.008.8249.0911.33
IDEFICS3-LLaMA3-8BOpen-Source
7.629.361.963.865.5737.2710.66
Ovis1.6-Gemma2-9BOpen-Source
1.328.480.281.937.8943.9410.56
IDEFICS2-8BOpen-Source
2.324.391.401.618.1235.458.87

Contact

We welcome feedback, suggestions, and questions about DocHop. If you encounter any issues with the benchmark or are interested in collaboration, feel free to reach out at zhuoran.yu@wisc.edu.