2 citations · 7 across the 15 of their papers we have counts for
6 papers · 1 filter
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
Yiyang Ma, Feng Zhou, Xuedan Yin +3
Leveraging pre-trained Diffusion Transformers (DiTs) for high-resolution (HR) image synthesis often leads to spatial layout collapse and degraded texture fidelity. Prior work mitig…
MaskSem: Semantic-Guided Masking for Learning 3D Hybrid High-Order Motion Representation
Wei Wei, Shaojie Zhang, Yonghao Dang +1
Human action recognition is a crucial task for intelligent robotics, particularly within the context of human-robot collaboration research. In self-supervised skeleton-based action…
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
Shaojie Zhang, Jiahui Yang, Jianqin Yin +2
Multimodal Large Language Models (MLLMs) have demonstrated significant success in visual understanding tasks. However, challenges persist in adapting these models for video compreh…
3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
Yixun Zhang, Lizhi Wang, Junjun Zhao +4
Camera-based perception in connected and autonomous vehicles remains exposed to physical adversarial attacks. Prior attacks often either optimize image-plane textures, weakening cr…
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
Huilai Li, Yonghao Dang, Ying Xing +2
Dense audio-visual event localization (DAVE) aims to identify event categories and locate the temporal boundaries in untrimmed videos. Most studies only employ event-related semant…
GaussianCAD: Robust Self-Supervised CAD Reconstruction from Three Orthographic Views Using 3D Gaussian Splatting
Zheng Zhou, Zhe Li, Bo Yu +8
The automatic reconstruction of 3D computer-aided design (CAD) models from CAD sketches has recently gained significant attention in the computer vision community. Most existing me…