activity
20242026
collaborators

6 papers

cs.CV2026

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

Jiesong Lian, Zixiang Zhou, Ruizhe Zhong +6

Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Repres…

cs.CV2025

Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation

Qiao Yu, Xianzhi Li, Yuan Tang +4

Generating 3D meshes from a single image is an important but ill-posed task. Existing methods mainly adopt 2D multiview diffusion models to generate intermediate multiview images,…

cs.CV2025

PointDreamer: Zero-shot 3D Textured Mesh Reconstruction from Colored Point Cloud

Qiao Yu, Xianzhi Li, Yuan Tang +4

Faithfully reconstructing textured meshes is crucial for many applications. Compared to text or image modalities, leveraging 3D colored point clouds as input (colored-PC-to-mesh) o…

cs.CV2025

More Text, Less Point: Towards 3D Data-Efficient Point-Language Understanding

Yuan Tang, Xu Han, Xianzhi Li +5

Enabling Large Language Models (LLMs) to comprehend the 3D physical world remains a significant challenge. Due to the lack of large-scale 3D-text pair datasets, the success of LLMs…

cs.CL2024

Natural Language Fine-Tuning

Jia Liu, Yue Wang, Zhiqi Lin +3

Large language model fine-tuning techniques typically depend on extensive labeled data, external guidance, and feedback, such as human alignment, scalar rewards, and demonstration.…

cs.AI2024

DCMAC: Demand-aware Customized Multi-Agent Communication via Upper Bound Training

Dongkun Huo, Huateng Zhang, Yixue Hao +4

Efficient communication can enhance the overall performance of collaborative multi-agent reinforcement learning. A common approach is to share observations through full communicati…