collaborators

17 papers

cs.CV2026

Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation

Jiabing Yang, Yixiang Chen, Yuan Xu +6

Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track o…

cs.RO2026

Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision

Yuan Xu, Yixiang Chen, Kai Wang +5

Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all t…

cs.CV2026

When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection

Tao Yu, Yujia Yang, Shenghua Chai +17

Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spliced across sources, or augme…

cs.CL2026

LaPA: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation

Jiabing Yang, Yixiang Chen, Zichen Wen +8

Prefix-based methods have emerged as a promising paradigm for Controllable Text Generation (CTG) due to their parameter efficiency. However, while effective in short sequences, the…

cs.CL2026

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization

Yan Sun, Guoxia Wang, Jinle Zeng +6

Pretraining large language models (LLMs) with next-token prediction has led to remarkable advances, yet the context-dependent nature of token embeddings in such models results in h…

cs.SD2026

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

Tao Yu, yiming ding, Shenghua Chai +16

Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively…