3 papers
cs.CV2026
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
Yian Li, Yang Jiao, Bin Zhu +4
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language…
cs.CL2026
Enhancing Action and Ingredient Modeling for Semantically Grounded Recipe Generation
Guoshan Liu, Bin Zhu, Yian Li +3
Recent advances in Multimodal Large Language Models (MLMMs) have enabled recipe generation from food images, yet outputs often contain semantically incorrect actions or ingredients…
cs.CV2025
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning
Yian Li, Wentao Tian, Yang Jiao +5
Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and ext…