instruction tuning 1large language models 1on-policy distillation 1reasoning 1reinforcement learning 1
From the 1 of 38 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
VisualScratchpad: Inference-time Visual Concepts Analysis in Vision Language Models
Hyesu Lim, Jinho Choi, Taekyung Kim +3
High-performing vision language models still produce incorrect answers, yet their failure modes are often difficult to explain. To make model internals more accessible and enable s…
cs.AI2025
What Defines Good Reasoning in LLMs? Dissecting Reasoning Steps with Multi-Aspect Evaluation
Heejin Do, Jaehui Hwang, Dongyoon Han +2
Evaluating large language models (LLMs) on final-answer correctness is the dominant paradigm. This approach, however, provides a coarse signal for model improvement and overlooks t…