collaborators

9 papers

cs.CV2026

Conversational Human Audio-visual Talking Dialogue Generation

Junhao Song, Lluis Guasch, Xilin He +8

Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. Howev…

cs.LG2026

SLOPE: Optimistic Potential Landscape Shaping for Model-based Reinforcement Learning

Yao-Hui Li, Zeyu Wang, Xin Li +7

Model-based reinforcement learning (MBRL) is sample-efficient but struggles in sparse reward settings. A critical bottleneck arises from the lack of informative gradients in sparse…

cs.CV2026

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration

Zhongyu Yang, Zuhao Yang, Shuo Zhan +3

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. A…

cs.IR2026

XR: Cross-Modal Agents for Composed Image Retrieval

Zhongyu Yang, Wei Pang, Yingfang Yuan

Retrieval is being redefined by agentic AI, demanding multimodal reasoning beyond conventional similarity-based paradigms. Composed Image Retrieval (CIR) exemplifies this shift as…

cs.AI2026

STProtein: predicting spatial protein expression from multi-omics data

Zhaorui Jiang, Yingfang Yuan, Lei Hu +1

The integration of spatial multi-omics data from single tissues is crucial for advancing biological research. However, a significant data imbalance impedes progress: while spatial…

cs.CV2025

InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration

Zhongyu Yang, Yingfang Yuan, Xuanming Jiang +2

Hallucination remains a critical challenge in large language models (LLMs), hindering the development of reliable multimodal LLMs (MLLMs). Existing solutions often rely on human in…