Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
PARM: Pipeline-Adapted Reward Model
Xingyu Fan, Wei Shao, Jiacheng Liu +2
Reward models (RMs) are central to aligning large language models (LLMs) with human preferences, powering RLHF and advanced decoding strategies. While most prior work focuses on si…
cs.AI2026
Not Search, But Scan: Benchmarking MLLMs on Scan-Oriented Academic Paper Reasoning
Rongjin Li, Zichen Tang, Xianghe Wang +9
With the rapid progress of multimodal large language models (MLLMs), AI already performs well at literature retrieval and certain reasoning tasks, serving as a capable assistant to…