7 papers
Hyperbolic Hierarchical Clustering for Visual Representation Learning
Jianan Wei, Guikun Chen, Zhiyuan Weng +3
We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An effective token mixer is a fundamental compo…
CoVStream: Edge-Cloud Collaboration for Understanding of Long Video Streams
Xu Liu, Guikun Chen, Zihao Yan +2
Long, continuous video streams are an increasingly critical driver of multimedia intelligence. Existing efforts often handle long videos with a sample-encode-reason approach using…
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
Mu Chen, Liulei Li, Wenguan Wang +1
Top-leading solutions for Video Scene Graph Generation (VSGG) typically adopt an offline pipeline. Though demonstrating promising performance, they remain unable to handle real-tim…
A Survey of World Models for Autonomous Driving
Tuo Feng, Wenguan Wang, Yi Yang
Recent breakthroughs in autonomous driving have been propelled by advances in robust world modeling, fundamentally transforming how vehicles interpret dynamic scenes and execute sa…
Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models
Liulei Li, Wenguan Wang, Yi Yang
Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though…
Vision-Language Navigation with Energy-Based Policy
Rui Liu, Wenguan Wang, Yi Yang
Vision-language navigation (VLN) requires an agent to execute actions following human instructions. Existing VLN models are optimized through expert demonstrations by supervised be…