VisWordBench: Bridging the Gap in Cross-modal Reasoning for Multimodal Large Language Models
Abstract
Although recent multimodal large language models (MLLMs)have advanced rapidly in vision–language reasoning, their capability bound-aries, particularly in cross-modal integration and reasoning, remain un-derexplored. Existing benchmarks primarily focus on evaluating uni-modal or loosely coupled multimodal abilities, leaving a gap in assess-ing complex cross-modal reasoning. To address this gap, we introduceVisWordBench, a benchmark comprising 2,625 English and 2,000 Chi-nese visual word puzzles with detailed human annotations. The bench-mark is designed to evaluate not only fundamental perceptual under-standing and world knowledge but also deep cross-modal reasoning skillsthat require consistently integrating visual and linguistic information.Through evaluations on VisWordBench, we find that current MLLMs ex-hibit under-diversified hypothesis search during reasoning. We thereforepropose a data construction pipeline and a training-free inference-timesteering strategy that promotes more diverse hypothesis exploration dur-ing CoT reasoning. Experimental results show consistent improvementsin reasoning quality and hypothesis diversity, supporting the utility ofVisWordBench for studying multimodal reasoning. The benchmark isavailable at: https://zenodo.org/records/20957955.