1 paper
Qiman Wu, Hanlin Chen, Lyujie Chen +19
Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, temporal modeling, and VLM integra…