activity
20182026
most citedLarge Language Model as Attributed Training Data Generator: A Tale of Diversity and Bias

72 citations · 180 across the 43 of their papers we have counts for

collaborators

48 papers

cs.AI2026

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering

Rushi Qiang, Changhao Li, Haotian Sun +3

Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expensive and feedback-driven environment inte…

cs.IR2026

RL-Index: Reinforcement Learning for Retrieval Index Reasoning

Yongjia Lei, Nedim Lipka, Zhisheng Qi +7

Retrieving external knowledge is crucial for real-world tasks but remains difficult when queries and relevant knowledge are linked by implicit reasoning (e.g., shared theorems or c…

cs.LG2026

Exploration-Driven Optimization for Test-Time Large Language Model Reasoning

Changhao Li, Yuchen Zhuang, Chenxiao Gao +4

Post-training techniques combined with inference-time scaling significantly enhance the reasoning and alignment capabilities of large language models (LLMs). However, a fundamental…

cs.CV2026

MEDVISTAGYM: A Scalable Training Environment for Thinking with Medical Images via Tool-Integrated Reinforcement Learning

Meng Lu, Yuxing Lu, Yuchen Zhuang +6

Vision language models (VLMs) achieve strong performance on general image understanding but struggle to think with medical images, especially when performing multi-step reasoning t…

cs.AI2025

Towards a Science of Scaling Agent Systems

Yubin Kim, Ken Gu, Chanwoo Park +17

Agents, language model-based systems capable of reasoning, planning, and acting are widely adopted in real-world tasks, yet how their performance changes as these systems scale acr…

cs.AI2025

Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs

Meng Lu, Ran Xu, Yi Fang +14

While recent vision-language models (VLMs) demonstrate strong image understanding, their ability to "think with images", i.e., to reason through multi-step visual interactions, rem…