Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models
Rituraj Sharma, Tu Vu
Looped language models turn hidden states into runtime state: each state is decoded for prediction and fed back into future computation. This creates a basic supervision question:…
cs.LG2026
The Master Key Hypothesis: Unlocking Cross-Model Capability Transfer via Linear Subspace Alignment
Rishab Balasubramanian, Pin-Jie Lin, Rituraj Sharma +6
We investigate whether post-trained capabilities can be transferred across models without retraining, with a focus on transfer across different model scales. We propose the Master…