14 papers
Conversational Human Audio-visual Talking Dialogue Generation
Junhao Song, Lluis Guasch, Xilin He +8
Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. Howev…
Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
Xuqing Yang, Yi Yuan, Shanzhe Lei +1
Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing R…
SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning
Yao-Hui Li, Zeyu Wang, Xin Li +7
Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse…
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration
Yi Yuan, Xuhong Wang, Shanzhe Lei
As agent-based systems continue to evolve, deep research agents are capable of automatically generating research-style reports across diverse domains. While these agents promise to…
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
Zhongyu Yang, Zuhao Yang, Shuo Zhan +3
Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. A…
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
Dannong Xu, Zhongyu Yang, Jun Chen +6
Multimodal large language models (MLLMs) achieve strong performance on benchmarks that evaluate text, image, or video understanding separately. However, these settings do not asses…