17 papers
Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation
Jiabing Yang, Yixiang Chen, Yuan Xu +6
Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track o…
Improving Vision-Language-Action Model Fine-Tuning with Structured Stage and Keyframe Supervision
Yuan Xu, Yixiang Chen, Kai Wang +5
Vision-Language-Action (VLA) models have shown strong potential for generalizable robotic manipulation. During fine-tuning, however, action supervision applies equally across all t…
When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection
Tao Yu, Yujia Yang, Shenghua Chai +17
Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spliced across sources, or augme…
LaPA: Length-Aware Prefix and Prompt Attention Augmentation for Long-Form Controllable Text Generation
Jiabing Yang, Yixiang Chen, Zichen Wen +8
Prefix-based methods have emerged as a promising paradigm for Controllable Text Generation (CTG) due to their parameter efficiency. However, while effective in short sequences, the…
SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization
Yan Sun, Guoxia Wang, Jinle Zeng +6
Pretraining large language models (LLMs) with next-token prediction has led to remarkable advances, yet the context-dependent nature of token embeddings in such models results in h…
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
Tao Yu, yiming ding, Shenghua Chai +16
Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start from audio alone and actively…