activity
20242026
collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2025

InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs

Bin Li, Ruichi Zhang, Han Liang +6

Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overl…

cs.CV2025

Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis

Kaiyang Ji, Ye Shi, Zichen Jin +5

Real-time synthesis of physically plausible human interactions remains a critical challenge for immersive VR/AR systems and humanoid robotics. While existing methods demonstrate pr…

cs.CV2025

MouseGPT: A Large-scale Vision-Language Model for Mouse Behavior Analysis

Teng Xu, Taotao Zhou, Youjia Wang +12

Analyzing animal behavior is crucial in advancing neuroscience, yet quantifying and deciphering its intricate dynamics remains a significant challenge. Traditional machine vision a…

cs.CV2025

UniDB: A Unified Diffusion Bridge Framework via Stochastic Optimal Control

Kaizhen Zhu, Mokai Pan, Yuexin Ma +4

Recent advances in diffusion bridge models leverage Doob's -transform to establish fixed endpoints between distributions, demonstrating promising results in image translation an…

cs.CV2024

SeqAfford: Sequential 3D Affordance Reasoning via Multimodal Large Language Model

Chunlin Yu, Hanqing Wang, Ye Shi +4

3D affordance segmentation aims to link human instructions to touchable regions of 3D objects for embodied manipulations. Existing efforts typically adhere to single-object, single…

cs.CV2024

NLPrompt: Noise-Label Prompt Learning for Vision-Language Models

Bikang Pan, Qun Li, Xiaoying Tang +6

The emergence of vision-language foundation models, such as CLIP, has revolutionized image-text representation, enabling a broad range of applications via prompt learning. Despite…