3 papers
cs.CV2026
Reasoning Matters for 3D Visual Grounding
Hsiang-Wei Huang, Kuang-Ming Chen, Wenhao Chai +3
The recent development of Large Language Models (LLMs) with strong reasoning ability has driven research in various domains such as mathematics, coding, and scientific discovery. M…
cs.CV2025
Warehouse Spatial Question Answering with LLM Agent
Hsiang-Wei Huang, Jen-Hao Cheng, Kuang-Ming Chen +8
Spatial understanding has been a challenging task for existing Multi-modal Large Language Models~(MLLMs). Previous methods leverage large-scale MLLM finetuning to enhance MLLM's sp…
cs.CV2025
TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
Jen-Hao Cheng, Vivian Wang, Huayu Wang +11
Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models. Existing methods either compress vid…