activity
20242026
collaborators

7 papers

cs.AI2026

PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents

Zhuohan Gu, Qizheng Zhang, Omar Khattab +1

Large language model (LLM) agents increasingly operate over long and recurring external contexts, like document corpora and code repositories. Across invocations, existing approach…

cs.LG2026

Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls

Shubhangi Upasani, Chen Wu, Jay Rainton +4

Test-time adaptation enables large language models (LLMs) to modify their behavior at inference without updating model parameters. A common approach is many-shot prompting, where l…

cs.LG2026

An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence

Qizhen Zhang, Ankush Garg, Jakob Foerster +3

Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulat…

cs.CL2025

BTS: Harmonizing Specialized Experts into a Generalist LLM

Qizhen Zhang, Prajjwal Bhargava, Chloe Bi +9

We present Branch-Train-Stitch (BTS), an efficient and flexible training algorithm for combining independently trained large language model (LLM) experts into a single, capable gen…

cs.CL2024

Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts

Nikolas Gritsch, Qizhen Zhang, Acyr Locatelli +2

Efficiency, specialization, and adaptability to new data distributions are qualities that are hard to combine in current Large Language Models. The Mixture of Experts (MoE) archite…

cs.LG2024

BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts

Qizhen Zhang, Nikolas Gritsch, Dwaraknath Gnaneshwar +8

The Mixture of Experts (MoE) framework has become a popular architecture for large language models due to its superior performance over dense models. However, training MoEs from sc…