collaborators

13 papers

cs.CV2026

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation

Jiabing Yang, Yixiang Chen, Yuan Xu +6

Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track o…

cs.RO2026

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Yixiang Chen, Peiyan Li, Yuan Xu +13

World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for co…

cs.RO2026

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

Yuan Xu, Yixiang Chen, Kai Wang +5

Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all t…

cs.CV2026

When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection

Tao Yu, Yujia Yang, Shenghua Chai +17

Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spliced across sources, or augme…

cs.CV2026

ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search

Tao Yu, Haopeng Jin, Hao Wang +18

In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Ope…

cs.DL2026

PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG

Tao Yu, Minghui Zhang, Zhiqing Cui +17

Transforming scientific papers into multimodal presentation content is essential for research dissemination but remains labor intensive. Existing automated solutions typically trea…