5 papers
Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization
Xuanyu Zhu, Yan Bai, Yang Shi +4
Representation autoencoders that reuse frozen pretrained vision encoders as visual tokenizers have achieved strong reconstruction and generation quality. However, existing methods…
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
Xiran Zhao, Jing Jin, Yan Bai +6
Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which do not fully reflect continuo…
Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization
Yi Zhang, Che Liu, Xiancong Ren +17
Developing a universal and versatile embodied intelligence system presents two primary challenges: the critical embodied data bottleneck, where real-world data is scarce and expens…
GLRD: Global-Local Collaborative Reason and Debate with PSL for 3D Open-Vocabulary Detection
Xingyu Peng, Si Liu, Chen Gao +4
The task of LiDAR-based 3D Open-Vocabulary Detection (3D OVD) requires the detector to learn to detect novel objects from point clouds without off-the-shelf training labels. Previo…
Training Video Foundation Models with NVIDIA NeMo
Zeeshan Patel, Ethan He, Parth Mannan +26
Video Foundation Models (VFMs) have recently been used to simulate the real world to train physical AI systems and develop creative visual experiences. However, there are significa…