3 papers
cs.CL2026
Doc-to-Atom: Learning to Compile and Compose Memory Atoms
Xingjian Diao, Wenbo Li, Yashas Malur Saidutta +3
Long input sequences are central to document understanding and multi-step reasoning in Large Language Models, yet the quadratic cost of attention makes inference both memory-intens…
cs.CV2026
Foveated Reasoning: Stateful, Action-based Visual Focusing for Vision-Language Models
Juhong Min, Lazar Valkov, Vitali Petsiuk +2
Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this tension via foveation: a coarse…
cs.CV2026
Towards High-Fidelity Gaussian Splatting with Queried-Convolution Neural Networks
Abhinav Kumar, Tristan Aumentado-Armstrong, Lazar Valkov +4
Gaussian Splatting has revolutionized the field of Novel View Synthesis (NVS) with faster training and real-time rendering. However, its reconstruction fidelity still trails behind…