activity
20242026
most citedMarco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models

1 citations · 2 across the 42 of their papers we have counts for

collaborators
Showing cs.CLShow all

29 papers · 1 filter

cs.CL2026

Neuron-Guided Fine-Tuning: Unlocking Efficient Alignment Mechanisms for Large Language Models

Zeyu Wu, Junchao Wu, Shudong Liu +8

Existing Supervised Fine-Tuning paradigms, particularly Full Parameter Fine-Tuning are often plagued by parameter redundancy, inconsistent data quality, and catastrophic forgetting…

cs.CL2026

UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation Paradigms

Peng Lai, Yichao Du, Junchao Wu +6

Reinforcement learning (RL) excels on tasks with verifiable rewards, but in open-ended tasks, the reliability of reward models remains a key challenge. Existing solutions either de…

cs.CL2026

WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs

Yiming Yao, Chenyang Lyu, Xuanfan Ni +4

Long-form audio inputs make the KV cache the dominant memory cost of speech LLMs. Prefill-only KV compression methods permanently discard audio KV positions once evicted, with no p…

cs.CL2026

STAR : Sentence Translation Alignment Rate for Document-to-Document Machine Translation

Yichen Dong, Hao Wang, Junhui Li +3

Large Language Models (LLMs) have enabled a shift from sentence-level to document-to-document (Doc2Doc) machine translation, promising improved global coherence. However, document-…

cs.CL2026

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment

Tianyu Dong, Yangyang Liu, Jiang Zhou +9

Sparse Mixture-of-Experts (MoE) architectures have emerged as an increasingly influential paradigm as they offer a strategic balance between parameter scalability and computational…

cs.CL2026

StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs

Youxin Zhu, Yixuan Ding, Peng Lai +3

Statistical analysis is a broad, complex field requiring both domain knowledge and tool proficiency. While prior work has evaluated large language models (LLMs) in this domain, exi…