2 papers
cs.AI2026
Entropy-Guided Data-Efficient Training for Multimodal Reasoning Reward Models
Shidong Yang, Tongwen Huang, Hao Wen +3
Multimodal reward models are crucial for aligning multimodal large language models with human preferences. Recent works have incorporated reasoning capabilities into these models,…
cs.IR2026
Real-Time Trend Prediction via Continually-Aligned LLM Query Generation
Zijing Hui, Wenhan Lyu, Shusen Wang +2
Trending news detection in low-traffic search environments faces a fundamental cold-start problem, where a lack of query volume prevents systems from identifying emerging or long-t…