activity
20232026
most citedLoRAMoE: Alleviate World Knowledge Forgetting in Large Language Models via MoE-Style Plugin

8 citations · 15 across the 61 of their papers we have counts for

collaborators

67 papers

cs.AI2026

ReDraft, Don't Just Distill: Reference-Driven Revision for Continual VLLM Post-Training

Zhihao Zhang, Mingqi Wu, Qiaole Dong +15

Continual post-training of large multimodal models should add new capabilities while preserving those from pre-training, and the two goals pull in opposite directions. SFT gives ex…

cs.AI2026

Atria Dawn: The Dawn of Agentic Superintelligence

Honglin Guo, Tao Gui, Yicheng Chen +139

As AI agents become participants in the development of their successors, they reshape both the production of intelligence and the role of human researchers. We introduce Atria Dawn…

cs.AI2026

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Jiaqiang Li, Yajie Yang, Zhiheng Xi +15

Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-s…

cs.LG2026

A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation

Bing Shao, Jiazheng Zhang, Long Ma +11

On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens rem…

cs.AI2026

CAFE: Self-Improving Search Agents Need Co-Evolving Feedback

Boyang Liu, Senjie Jin, Peixin Wang +17

Reliable search requires more than acquiring external evidence. An agent must also recognize and recover from errors as its trajectory unfolds. In-trajectory feedback provides a me…

cs.CL2026

IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations

Dingwei Zhu, Jiahan Li, Chengjun Pan +22

Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustness via implicit history sca…