14 papers
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
Jacob Dineen, Aswin RRV, Zhikun Xu +1
Co-evolutionary self-play, where one language model generates problems and another solves them, promises curriculum learning without human supervision. The promise breaks down earl…
Skill Reuse as Compression in Agentic RL
Zhikun Xu, Yu Feng, Jacob Dineen +3
Large language model agents trained with reinforcement learning (RL) often learn brittle, task-specific shortcuts. We hypothesize that agents generalize better when their successfu…
VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
Zhaonan Li, Kyle R. Chickering, Bangzheng Li +13
A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties…
CORE: Concept-Oriented Reinforcement for Bridging the Definition-Application Gap in Mathematical Reasoning
Zijun Gao, Zhikun Xu, Xiao Ye +1
Large language models (LLMs) often solve challenging math exercises yet fail to apply the concept right when the problem requires genuine understanding. Popular Reinforcement Learn…
Reliable Use of Lemmas via Eligibility Reasoning and SectionAware Reinforcement Learning
Zhikun Xu, Xiaodong Yu, Ben Zhou +6
Recent large language models (LLMs) perform strongly on mathematical benchmarks yet often misapply lemmas, importing conclusions without validating assumptions. We formalize lemma$…
Unbiased Visual Reasoning with Controlled Visual Inputs
Zhaonan Li, Shijie Lu, Fei Wang +11
End-to-end Vision-language Models (VLMs) often answer visual questions by exploiting spurious correlations instead of causal visual evidence, and can become more shortcut-prone whe…