1 paper · 1 filter
Qiman Wu, Hanlin Chen, Lyujie Chen +19
Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, temporal modeling, and VLM integra…