8 papers
Offline-Online Curriculum RL for Multimodal Reasoning
Wendi Deng, Hang Du, Guoshun Nan +11
Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers. This behavior undermines…
Uncertainty as Remedy: Mitigating Satisfaction Label Bias in Short Video Multi-Objective Ensemble Ranking
Zonghe Shao, Tiantian He, Xiaoxiao Xu +6
The core objective of short video recommendation is to model users' unobservable true satisfaction with recommended videos. As the dominant industrial framework, end-to-end multi-o…
SentGuard: Sentence-Level Streaming Guardrails for Large Language Models
Jiaqi Yu, Xin Wang, Yixu Wang +4
Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate. Existing guardrails fall int…
S2Aligner: Pair-Efficient and Transferable Pre-Training for Sparse Text-Attributed Graphs
Yuhan Wang, Haopeng Zhang, Yibo Ding +6
Pre-training on text-attributed graphs (TAGs) is central to building transferable graph foundation models, where LLM-as-Aligner methods align graph and text representations through…
TAME: Test-Time Adversarial Prompt Tuning via Mixture-of-Experts for Vision-Language Models
Xin Wang, Yixu Wang, Jiaming Zhang +6
Large-scale pre-trained Vision-Language models (VLMs), such as CLIP, exhibit strong zero-shot generalization, yet remain highly vulnerable to imperceptible adversarial perturbation…
LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs
Jingfeng Chen, Jiawen Qian, Wendi Deng +5
Video understanding in multimodal large language models requires selecting informative frames from long, redundant videos under limited visual-token budgets. Existing methods often…