most citedCURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.LG2026

Not All Negative Samples Are Equal: LLMs Learn Better from Plausible Reasoning

Zixiang Di, Jinyi Han, Shuo Zhang +8

Learning from negative samples holds great promise for improving Large Language Model (LLM) reasoning capability, yet existing methods treat all incorrect responses as equally info…

cs.LG2026

From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation

Yongqi Wang, Xiaofeng Ji, Jie Wang +6

Adapting Large Language Models (LLMs) to specialized domains without human-annotated data is a crucial yet formidable challenge. Widely adopted knowledge distillation methods often…

cs.CV2025

Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition

Zichen Liang, Jingjing Fei, Jie Wang +6

Recent advances in multimodal large language models (MLLMs) have been primarily evaluated on general-purpose benchmarks, while their applications in domain-specific scenarios, such…

cs.LG20251 cited

CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention

Qingbin Li, Rongkun Xue, Jie Wang +8

Recent advances in Reinforcement Learning with Verified Reward (RLVR) have driven the emergence of more sophisticated cognitive behaviors in large language models (LLMs), thereby e…

cs.CL2025

Mitigating Hallucinations in Large Language Models via Causal Reasoning

Yuangang Li, Yiqing Shen, Yi Nian +7

Large language models (LLMs) exhibit logically inconsistent hallucinations that appear coherent yet violate reasoning principles, with recent research suggesting an inverse relatio…

cs.CL2025

LeTS: Learning to Think-and-Search via Process-and-Outcome Reward Hybridization

Qi Zhang, Shouqing Yang, Lirong Gao +8

Large language models (LLMs) have demonstrated impressive capabilities in reasoning with the emergence of reasoning models like OpenAI-o1 and DeepSeek-R1. Recent research focuses o…