activity
20242026
most citedArgus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

Yifan Xu, Chao Zhang, Ruifei Ma +4

The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-lev…

cs.CV2025

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models

Xiaoyan Wang, Zeju Li, Yifan Xu +5

New era has unlocked exciting possibilities for extending Large Language Models (LLMs) to tackle 3D vision-language tasks. However, most existing 3D multimodal LLMs (MLLMs) rely on…

cs.CV20251 cited

Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models

Yifan Xu, Chao Zhang, Hanqi Jiang +6

Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend…

cs.CV2025

MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation

Zhifei Yang, Keyang Lu, Chao Zhang +9

Controllable 3D scene generation has extensive applications in virtual reality and interior design, where the generated scenes should exhibit high levels of realism and controllabi…

cs.CV2024

3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding

Zeju Li, Chao Zhang, Xiaoyan Wang +4

The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3…