14 citations · 29 across the 14 of their papers we have counts for
20 papers
Conversational Human Audio-visual Talking Dialogue Generation
Junhao Song, Lluis Guasch, Xilin He +8
Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. Howev…
Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
Xuqing Yang, Yi Yuan, Shanzhe Lei +1
Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and question-answering tasks. However, prevailing R…
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration
Yi Yuan, Xuhong Wang, Shanzhe Lei
As agent-based systems continue to evolve, deep research agents are capable of automatically generating research-style reports across diverse domains. While these agents promise to…
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration
Zhongyu Yang, Zuhao Yang, Shuo Zhan +3
Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. A…
MultiHaystack: Benchmarking Multimodal Retrieval and Reasoning over 40K Images, Videos, and Documents
Dannong Xu, Zhongyu Yang, Jun Chen +6
Multimodal large language models (MLLMs) achieve strong performance on benchmarks that evaluate text, image, or video understanding separately. However, these settings do not asses…
XR: Cross-Modal Agents for Composed Image Retrieval
Zhongyu Yang, Wei Pang, Yingfang Yuan
Retrieval is being redefined by agentic AI, demanding multimodal reasoning beyond conventional similarity-based paradigms. Composed Image Retrieval (CIR) exemplifies this shift as…