13 papers
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Lai Wei, Chengqi Li, Jiapeng Li +3
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as co…
Distilled Reinforcement Learning for LLM Post-training
Chen Wang, Zhaochun Li, Jionghao Bai +4
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL)…
Belief-Aware VLM Model for Human-like Reasoning
Anshul Nayak, Shahil Shaik, Yue Wang
Traditional neural network models for intent inference rely heavily on observable states and struggle to generalize across diverse tasks and dynamic environments. Recent advances i…
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
Ruina Hu, Chen Wang, Lai Wei +5
Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers. However, such outcome-only re…
HTAM: Hierarchical Transition-Attended Memory for Operator Optimization
Yining Zhang, Mingyang Yi, Chen Wang +5
High-performance GPU kernels are essential for efficient LLM deployment, yet optimizing them remains expertise-intensive. Recent LLM-based code generation makes automatic GPU opera…
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
Chen Wang, Zhaochun Li, Jionghao Bai +3
Reinforcement learning (RL) is a key paradigm for post-training large language models (LLMs), but the widely used Group Relative Policy Optimization (GRPO) often suffers from entro…