1 citations · 1 across the 6 of their papers we have counts for
6 papers · 1 filter
S-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
Beining Xu, Siting Zhu, Zhao Jin +2
3D Visual Grounding (3DVG) focuses on locating objects in 3D scenes based on natural language descriptions, serving as a fundamental task for embodied AI and robotics. Recent advan…
Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
Junxian Li, Xinyue Xu, Sai Ma +2
Multimodal Large Language Models (MLLMs) frequently suffer from unfaithfulness, generating reasoning chains that drift from visual evidence or contradict final predictions. We prop…
AI for Service: Proactive Assistance with AI Glasses
Zichen Wen, Yiyu Wang, Chenfei Liao +10
In an era where AI is evolving from a passive tool into an active and adaptive companion, we introduce AI for Service (AI4Service), a new paradigm that enables proactive and real-t…
MusicScaffold: Bridging Machine Efficiency and Human Growth in Adolescent Creative Education through Generative AI
Zhejing Hu, Yan Liu, Zhi Zhang +4
Adolescence is marked by strong creative impulses but limited strategies for structured expression, often leading to frustration or disengagement. While generative AI lowers techni…
Mol-R1: Towards Explicit Long-CoT Reasoning in Molecule Discovery
Jiatong Li, Weida Wang, Qinggang Zhang +6
Large language models (LLMs), especially Explicit Long Chain-of-Thought (CoT) reasoning models like DeepSeek-R1 and QWQ, have demonstrated powerful reasoning capabilities, achievin…
Control-R: Towards controllable test-time scaling
Di Zhang, Weida Wang, Junxian Li +10
This paper target in addressing the challenges of underthinking and overthinking in long chain-of-thought (CoT) reasoning for Large Reasoning Models (LRMs) by introducing Reasoning…