6 papers
SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models
Jiesong Lian, Zixiang Zhou, Ruizhe Zhong +6
Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Repres…
Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation
Qiao Yu, Xianzhi Li, Yuan Tang +4
Generating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images,…
PointDreamer: Zero-shot 3D Textured Mesh Reconstruction from Colored Point Cloud
Qiao Yu, Xianzhi Li, Yuan Tang +4
Faithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) o…
More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding
Yuan Tang, Xu Han, Xianzhi Li +5
Enabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs…
Natural Language Fine-Tuning
Jia Liu, Yue Wang, Zhiqi Lin +3
Large language model fine-tuning techniques typically depend on extensive labeled data, external guidance, and feedback, such as human alignment, scalar rewards, and demonstration.…
DCMAC: Demand-aware Customized Multi-Agent Communication via Upper Bound Training
Dongkun Huo, Huateng Zhang, Yixue Hao +4
Efficient communication can enhance the overall performance of collaborative multi-agent reinforcement learning. A common approach is to share observations through full communicati…