activity
20242026
collaborators

8 papers

cs.CV2026

Conversational Human Audio-visual Talking Dialogue Generation

Junhao Song, Lluis Guasch, Xilin He +8

Large-scale dyadic interactive audio-visual dialogue (DIAD) datasets provide fundamental data resources for developing humanoid interactive virtual agents and digital humans. Howev…

cs.CV2026

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning

Ankan Deria, Komal Kumar, Xilin He +4

Recent vision-language models (VLMs) typically rely on a single vision encoder trained with contrastive image-text objectives, such as CLIP-style pretraining. While contrastive enc…

cs.CV2025

SynFER: Towards Boosting Facial Expression Recognition with Synthetic Data

Xilin He, Cheng Luo, Xiaole Xian +8

Facial expression datasets remain limited in scale due to the subjectivity of annotations and the labor-intensive nature of data collection. This limitation poses a significant cha…

cs.CV2025

BOTM: Echocardiography Segmentation via Bi-directional Optimal Token Matching

Zhihua Liu, Lei Tong, Xilin He +4

Existed echocardiography segmentation methods often suffer from anatomical inconsistency challenge caused by shape variation, partial observation and region ambiguity with similar…

cs.CV2025

Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation

Zhihua Liu, Amrutha Saseendran, Lei Tong +8

Open-set image segmentation poses a significant challenge because existing methods often demand extensive training or fine-tuning and generally struggle to segment unified objects…

cs.LG2025

Benchmarking Graph Representations and Graph Neural Networks for Multivariate Time Series Classification

Wennuo Yang, Shiling Wu, Yuzhi Zhou +5

Multivariate Time Series Classification (MTSC) enables the analysis if complex temporal data, and thus serves as a cornerstone in various real-world applications, ranging from heal…