3 papers
cs.CV2026
BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
Shengao Wang, Wenqi Wang, Zecheng Wang +20
Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded…
cs.CV2026
A Training-Free Guess What Vision Language Model from Snippets to Open-Vocabulary Object Detection
Guiying Zhu, Bowen Yang, Yin Zhuang +5
Open-Vocabulary Object Detection (OVOD) aims to develop the capability to detect anything. Although myriads of large-scale pre-training efforts have built versatile foundation mode…
cs.AI2025
Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation
Ning Wang, Zihan Yan, Weiyang Li +3
Embodied agents exhibit immense potential across a multitude of domains, making the assurance of their behavioral safety a fundamental prerequisite for their widespread deployment.…