Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing
Haowen Gao, Haoran Chen, Can Wang +5
Agent skills are structured procedural packages that guide frozen LLM agents in specialized workflows. Skills rarely remain sufficient after deployment: edge cases, API changes, an…
cs.AI2026
DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage
Haowen Gao, Zhenyu Zhang, Liang Pang +7
Reinforcement learning (RL) with group relative policy optimization (GRPO) has become a widely adopted approach for enhancing the reasoning capabilities of multimodal large languag…