Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
Pengkai Wang, Pengwei Liu, Qi Zuo +3
Reinforcement learning (RL) has powered many recent breakthroughs in large language models (LLMs), especially for tasks where rewards can be computed automatically, such as code ge…
cs.CL2025
A Comprehensive FP8 Training Recipe for Reasoning-Enhanced Language Models
Wenjun Wang, Shuo Cai, Congkai Xie +7
The immense computational cost of training Large Language Models (LLMs) presents a major barrier to innovation. While FP8 training offers a promising solution with significant theo…
cs.CL2025
InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning
Zeyu Liu, Zhitian Hou, Guanghao Zhu +3
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in domains such as visual understanding and mathematical reasoning. However, their application in the med…