activity
20242026
collaborators

7 papers

cs.CL2026

Dynamic Rollout Editing for Reducing Overthinking in RL-Trained Reasoning Models

Zihao Wei, Wenjie Shi, Liang Pang +8

Long-form chain-of-thought reasoning can improve LLM performance on complex tasks, but models often continue generating unnecessary reasoning after a correct answer has emerged. We…

cs.AI2026

SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing

Haowen Gao, Haoran Chen, Can Wang +5

Agent skills are structured procedural packages that guide frozen LLM agents in specialized workflows. Skills rarely remain sufficient after deployment: edge cases, API changes, an…

cs.AI2026

ActiveMem: Distributed Active Memory for Long-Horizon LLM Reasoning

Yunhan Jiang, Wenbin Duan, Shasha Guo +3

Memory is essential for enabling large language model (LLM) agents to handle long-horizon reasoning tasks. Existing memory mechanisms are largely centralized, typically organizing…

cs.CV2026

GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation

Shasha Guo, Liang Pang, Xi Wang +3

Auxiliary lines are essential for solving complex geometric problems but remain challenging for large vision-language models (LVLMs). Recent attempts construct auxiliary lines via…

cs.CL2025

VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering

Yanling Wang, Yihan Zhao, Xiaodong Chen +7

Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA…

cs.CL2025

Diversifying Question Generation over Knowledge Base via External Natural Questions

Shasha Guo, Jing Zhang, Xirui Ke +2

Previous methods on knowledge base question generation (KBQG) primarily focus on enhancing the quality of a single generated question. Recognizing the remarkable paraphrasing abili…