works on

From the 1 of 28 linked papers with an AI index.

activity
20242026
collaborators

28 papers

cs.LG2026

HSD: Hybrid Hindsight Self-Distillation

Qiye Cai, Yichuan Ma, Linyang Li +7

Reinforcement learning with verifiable rewards (RLVR) provides reliable outcome supervision for language model reasoning, but a scalar trajectory reward offers limited token-level…

cs.CL2026

AI Can Learn Scientific Taste

Jingqi Tong, Mingzhe Li, Hangcheng Li +20

The paper introduces a reinforcement‑learning framework that uses citation‑based community feedback to train models that can judge the impact of scientific papers and generate high…

cs.CV2026

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang +14

The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual…

cs.CL2026

Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment

Yuming Yang, Mingyoung Lai, Wanxu Zhao +13

Long chain-of-thought (CoT) trajectories provide rich supervision signals for distilling reasoning from teacher to student LLMs. However, both prior work and our experiments show t…

cs.LG2026

What Makes Position Zero Special? A Mechanistic Study of Position Zero Attention Sinks in LLMs

Runyu Peng, Ruixiao Li, Mingshu Chen +5

Transformers frequently allocate disproportionate attention to specific tokens, a phenomenon known as attention sinks. Causal large language models reliably form one at position ze…

cs.LG2026

Explicit Multi-head Attention for Inter-head Interaction in Large Language Models

Runyu Peng, Yunhua Zhou, Demin Song +4

In large language models built upon the Transformer architecture, recent studies have shown that inter-head interaction can enhance attention performance. Motivated by this, we pro…