11 papers
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
Yunheng Li, Guohong Mu, Hao Li +4
Multimodal large language models (MLLMs) have become a prevailing paradigm for unified video perception. However, post-training on large multi-task datasets remains challenging, as…
RecoReward: Recommender-Guided Multimodal Description Generation for Recommendation
Guohong Mu, Yueyang Liu, Jiangxia Cao +8
Multimodal large language models (MLLMs) can convert multimodal item content into structured descriptions used as semantic features for recommendation. Conventional content-only ge…
Plan Before Search: Search Agents Need Plan
Zhipeng Qian, Zihan Liang, Yufei Ma +7
Training large language models as retrieval-augmented reasoning agents typically combines reinforcement learning with an SFT cold start distilled from a stronger model. However, th…
SLIP-RS: Structured-Attribute Language-Image Pre-Training for Remote Sensing Object Detection
Chenxu Wang, Yuxuan Li, Yunheng Li +3
Existing language-image pre-training for remote sensing object detection is constrained by Monolithic Label Learning, which relies on exhaustively enumerating open-set categories v…
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Boyuan Sun, Bowen Yin, Yuanming Li +2
We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts…
Degradation Frequency Curve: An Explicit Frequency-Quantified Representation for All-in-One Image Restoration
Xinghua Huang, Zhixiong Yang, Chen Wu +6
A fundamental difficulty in all-in-one blind image restoration is that degradation is usually treated as an implicit factor hidden in degraded-to-clean mapping, rather than as an e…