Skip to main navigation Skip to search Skip to main content

Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models

  • Zhongyu Yang
  • , Dannong Xu
  • , Yonghan Zhang
  • , Chen Kefan
  • , Xinyi Wang
  • , Xu Yang
  • , Wei Pang
  • , Yingfang Yuan*
  • *Corresponding author for this work

Research output: Contribution to journalConference articlepeer-review

Abstract

Unified Foundation Models (UFMs), which support interleaved multimodal generation and understanding, have been proposed as a promising paradigm for reasoning about dynamic world states, yet it remains unclear whether the visual content they generate functions as grounded evidence for subsequent reasoning or merely as auxiliary output. Existing benchmarks largely evaluate generation and understanding as separate capabilities and do not test their functional dependence during reasoning. We introduce \textbf{UFO}, a benchmark designed to evaluate whether UFMs generate and use image and text cues as evidence for compositional multimodal reasoning. UFO spans three cue types, state determination, state reconstruction, and state augmentation, which correspond to progressively smaller transformations of the underlying world state. Our analysis reveals a significant modality gap, as models often achieve high prediction accuracy even when the generated visual cues exert limited influence on their decisions, indicating weakened evidential coupling and a reliance on textual shortcuts rather than robust cross modal grounding.
Original languageEnglish
JournalProceedings of Machine Learning Research
Publication statusAccepted/In press - 27 May 2026
Event43rd International Conference on Machine Learning 2026 - Seoul, Korea, Republic of
Duration: 6 Jun 20266 Nov 2026
https://icml.cc/

Fingerprint

Dive into the research topics of 'Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models'. Together they form a unique fingerprint.

Cite this