6 papers
ObjectMorpher: 3D-Aware Image Editing via Deformable 3DGS Models
Yuhuan Xie, Aoxuan Pan, Yi-Hua Huang +4
Achieving precise, object-level control in image editing remains challenging: 2D methods lack 3D awareness and often yield ambiguous or implausible results, while existing 3D-aware…
Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language Models
Baoheng Zhang, Jiahui Liu, Gui Zhao +7
Multimodal Large Language Models (MLLMs) perform strong vision-language reasoning under standard conditions but fail in extreme illumination, where RGB inputs lose irrevocable stru…
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
Ouxiang Li, Yuan Wang, Xinting Hu +7
Text-to-image (T2I) generation aims to synthesize images from textual prompts, which jointly specify what must be shown and imply what can be inferred, which thus correspond to two…
ASSIST-3D: Adapted Scene Synthesis for Class-Agnostic 3D Instance Segmentation
Shengchao Zhou, Jiehong Lin, Jiahui Liu +3
Class-agnostic 3D instance segmentation tackles the challenging task of segmenting all object instances, including previously unseen ones, without semantic class reliance. Current…
Aligning Effective Tokens with Video Anomaly in Large Language Models
Yingxian Chen, Jiahui Liu, Ruidi Fan +6
Understanding abnormal events in videos is a vital and challenging task that has garnered significant attention in a wide range of applications. Although current video understandin…
NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding
Running Zhao, Zhihan Jiang, Xinchen Zhang +7
Users often take notes for instructional videos to access key knowledge later without revisiting long videos. Automated note generation tools enable users to obtain informative not…