collaborators

5 papers

cs.CV2026

Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models

Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai +4

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for…

cs.CV2026

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs

Jihwan Kim, Nikhil Parthasarathy, Danfeng Qin +5

The fundamental challenge in scaling Video Large Language Models (Video LLMs) to long-form video lies in managing the explosion of visual-token context length. Existing strategies…

cs.CV2026

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

Lee Hsin-Ying, Hanwen Jiang, Yiqun Mei +3

Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally incomplete. Such reliance often…

cs.CV2026

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

Deshui Miao, Xingsen Huang, Yameng Gu +3

Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearance. Recent multimodal models…

cs.CV2024

Chat-Edit-3D: Interactive 3D Scene Editing via Text Prompts

Shuangkang Fang, Yufeng Wang, Yi-Hsuan Tsai +4

Recent work on image content manipulation based on vision-language pre-training models has been effectively extended to text-driven 3D scene editing. However, existing schemes for…