From the 2 of 14 linked papers with an AI index.
4 papers · 1 filter
ATLAS: All-round Testing of Long-context Abilities across Scales
Deli Huang, Cunguang Wang, Hongyin Tang +15
Long-context language models now advertise context windows up to millions of tokens, yet evaluations typically report a single length or a narrow task family, masking two failure m…
Adaptive Test-Time Compute Allocation for Block Diffusion Language Models in Complex Reasoning
Yi Lu, Deyang Kong, Jianing Wang +8
Recent advances in block diffusion language models have demonstrated competitive performance and strong scalability on reasoning tasks. However, their test-time compute allocation…
LongCat-Flash Technical Report
Meituan LongCat Team, Bayan, Bei Li +179
We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming f…
Calibrating the Confidence of Large Language Models by Eliciting Fidelity
Mozhi Zhang, Mianqiu Huang, Rundong Shi +5
Large language models optimized with techniques like RLHF have achieved good alignment in being helpful and harmless. However, post-alignment, these language models often exhibit o…