6 papers
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
Rui Zhu, Xin Shen, Shuchen Wu +6
Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmark…
Decide Then Retrieve: A Training-Free Framework with Uncertainty-Guided Triggering and Dual-Path Retrieval
Wang Chen, Guanqiang Qi, Weikang Li +3
Retrieval-augmented generation (RAG) enhances large language models (LLMs) by incorporating external knowledge, but existing approaches indiscriminately trigger retrieval and rely…
FingerCap: Fine-grained Finger-level Hand Motion Captioning
Xin Shen, Rui Zhu, Lei Shen +10
Understanding fine-grained human hand motion is fundamental to visual perception, embodied intelligence, and multimodal communication. In this work, we propose Fine-grained Finger-…
Probabilistic Modeling of Intentions in Socially Intelligent LLM Agents
Feifan Xia, Yuyang Fang, Defang Li +5
We present a probabilistic intent modeling framework for large language model (LLM) agents in multi-turn social dialogue. The framework maintains a belief distribution over a partn…
Cross-LoRA: A Data-Free LoRA Transfer Framework across Heterogeneous LLMs
Feifan Xia, Mingyang Liao, Yuyang Fang +6
Traditional parameter-efficient fine-tuning (PEFT) methods such as LoRA are tightly coupled with the base model architecture, which constrains their applicability across heterogene…
PAIRS: Parametric-Verified Adaptive Information Retrieval and Selection for Efficient RAG
Wang Chen, Guanqiang Qi, Weikang Li +3
Retrieval-Augmented Generation (RAG) has become a cornerstone technique for enhancing large language models (LLMs) with external knowledge. However, current RAG systems face two cr…