Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
DARO: Difficulty-Aware Reweighting Policy Optimization
Jingyu Zhou, Lu Ma, Hao Liang +3
Recent advances in large language models (LLMs) have shown that reasoning ability can be significantly enhanced through Reinforcement Learning with Verifiable Rewards (RLVR). Group…
cs.CL2025
Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
Hao Liang, Ruitao Wu, Bohan Zeng +3
Multimodal reasoning remains a fundamental challenge in artificial intelligence. Despite substantial advances in text-based reasoning, even state-of-the-art models such as GPT-o3 s…