2 papers
cs.CV2026
StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs
Joya Chen, Zeyun Zhong, Mike Zheng Shou
Humans effortlessly perceive the present while remembering the past, yet streaming VLMs often trade off real-time perception against long-term memory. Prior work shows that shorten…
cs.CL2026
What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
Ziran Li, Qiang Wang, Zhengyu Chen +4
Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: co…