16 citations · 16 across the 13 of their papers we have counts for
15 papers · 1 filter
Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
Xiaogang Peng, Zeyu Han, Zichong Meng +4
Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most e…
Multi-Modal Decouple and Recouple Network for Robust 3D Object Detection
Rui Ding, Zhaonian Kuang, Yuzhe Ji +3
Multi-modal 3D object detection with bird's eye view (BEV) has achieved desired advances on benchmarks. Nonetheless, the accuracy may drop significantly in the real world due to da…
UniLayDiff: A Unified Diffusion Transformer for Content-Aware Layout Generation
Zeyang Liu, Le Wang, Sanping Zhou +4
Content-aware layout generation is a critical task in graphic design automation, focused on creating visually appealing arrangements of elements that seamlessly blend with a given…
SAMPO:Scale-wise Autoregression with Motion PrOmpt for generative world models
Sen Wang, Jingyi Tian, Le Wang +7
World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive…
LayoutRAG: Retrieval-Augmented Model for Content-agnostic Conditional Layout Generation
Yuxuan Wu, Le Wang, Sanping Zhou +3
Controllable layout generation aims to create plausible visual arrangements of element bounding boxes within a graphic design according to certain optional constraints, such as the…
Moment Quantization for Video Temporal Grounding
Xiaolong Sun, Le Wang, Sanping Zhou +5
Video temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishin…