14 papers
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Kailin Lyu, Di Wu, Pengwei Zhang +12
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…
PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning
Wenyuan Wang, Lianyu Hu, Hao Wang +1
Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unreliable in reasoning about physical plaus…
TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios
Kailin Lyu, Di Wu, Long Xiao +7
Among the five primary human senses, tactile is arguably the most fundamental to survival, as it enables the perception of physical contact and interaction in real-world environmen…
CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
Lianyu Hu, Shengqian Qin, Zeqin Liao +4
Chain-of-thought (CoT) reasoning has enabled multi-modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit intermediate reasoning steps…
TVI-CoT: Text-Visual Interleaved Chain-of-Thought Reasoning for Multimodal Understanding
Lianyu Hu, Xiaoyu Ma, Zeqin Liao +1
Chain-of-thought (CoT) reasoning has proven effective for enhancing problem-solving in large language models. However, when applied to multimodal LLMs (MLLMs), existing CoT approac…
PhotoCraft: Agentic Reasoning with Hierarchical Self-Evolving Memory for Deep Image Search
Kailin Lyu, Zhiqiang Yuan, Jianwei He +9
Deep Image Search requires multi-step reasoning over rich contextual cues, such as time, location, and event relations. However, most existing LLM-based agents are stateless and re…