Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models
Rituraj Sharma, Tu Vu
Looped language models turn hidden states into runtime state: each state is decoded for prediction and fed back into future computation. This creates a basic supervision question:…
cs.LG2024
What Matters for Model Merging at Scale?
Prateek Yadav, Tu Vu, Jonathan Lai +4
Model merging aims to combine multiple expert models into a more capable single model, offering benefits such as reduced storage and serving costs, improved generalization, and sup…