6 papers
OmniVLA-RL: A Vision-Language-Action Model with Spatial Understanding and Online RL
Haoxiang Jie, Yaoyuan Yan, Xiangyu Wei +4
Visual-Language-Action (VLA) models represent a paradigm shift in embodied AI, yet existing frameworks often struggle with imprecise spatial perception, suboptimal multimodal fusio…
One Dinomaly2 Detect Them All: A Unified Framework for Full-Spectrum Unsupervised Anomaly Detection
Jia Guo, Shuai Lu, Lei Fan +9
Unsupervised anomaly detection (UAD) has evolved from building specialized single-class models to unified multi-class models, yet existing multi-class models significantly underper…
Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact
Rizwan Qureshi, Ranjan Sapkota, Abbas Shah +16
Can machines truly think, reason and act in domains like humans? This enduring question continues to shape the pursuit of Artificial General Intelligence (AGI). Despite the growing…
YOLO advances to its genesis: a decadal and comprehensive review of the You Only Look Once (YOLO) series
Ranjan Sapkota, Marco Flores Calero, Rizwan Qureshi +9
This review systematically examines the progression of the You Only Look Once (YOLO) object detection algorithms from YOLOv1 to the recently unveiled YOLOv12. Employing a reverse c…
Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
Yikun Ji, Hong Yan, Jun Lan +5
The rapid advancement of image generation technologies intensifies the demand for interpretable and robust detection methods. Although existing approaches often attain high accurac…
Towards Scalable Topological Regularizers
Hiu-Tung Wong, Darrick Lee, Hong Yan
Latent space matching, which consists of matching distributions of features in latent space, is a crucial component for tasks such as adversarial attacks and defenses, domain adapt…