Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs
Yu Cheng, Arushi Goel, Hakan Bilen
Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlen…
cs.CV2025
Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study
Jihai Zhang, Tianle Li, Linjie Li +2
Unified vision-language models (VLMs) aim to support both visual understanding and generation within a single framework, but it remains unclear when mixed training benefits both ca…