activity
20242026
most citedSeeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

3 citations · 4 across the 8 of their papers we have counts for

collaborators

11 papers

cs.LG2026

Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning

Shaotian Yan, Kaiyuan Liu, Chen Shen +6

In this report, we introduce DASD-4B-Thinking, a lightweight yet highly capable, fully open-source reasoning model. It achieves SOTA performance among open-source models of compara…

cs.AI2025

SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning

Yue Xin, Chen Shen, Shaotian Yan +5

Chain-of-Thought (CoT) prompting enhances the math reasoning capability of large language models (LLMs) to a large margin. However, the mechanism underlying such improvements remai…

cs.CV2025

Enhancing Spatial Reasoning through Visual and Textual Thinking

Xun Liang, Xin Guo, Zhongming Jin +5

The spatial reasoning task aims to reason about the spatial relationships in 2D and 3D space, which is a fundamental capability for Visual Question Answering (VQA) and robotics. Al…

cs.CL2025

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning

Chenxi Huang, Shaotian Yan, Liang Xie +6

Representation Fine-tuning (ReFT), a recently proposed Parameter-Efficient Fine-Tuning (PEFT) method, has attracted widespread attention for significantly improving parameter effic…

cs.CL2025

Controlling Thinking Speed in Reasoning Models

Zhengkai Lin, Zhihang Fu, Ze Chen +6

Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at…

cs.CL2025

Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models

Kaiyuan Liu, Chen Shen, Zhanwei Zhang +3

While recent advances in large reasoning models have demonstrated remarkable performance, efficient reasoning remains critical due to the rapid growth of output length. Existing op…