21 citations · 39 across the 7 of their papers we have counts for
10 papers · 1 filter
CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios
Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G +1
While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpre…
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai +2
Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when co…
High-Quality Entity Segmentation and Grounding
Lu Qi, Yi-Wen Chen, Tao Zhang +4
In this work, we propose ESG, a pipeline for high-quality entity segmentation and grounding supported by a new dataset EntitySeg. At first, the proposed dataset naming EntitySeg co…
Text-Driven Image Editing via Learnable Regions
Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai +2
Language has emerged as a natural interface for image editing. In this paper, we introduce a method for region-based image editing driven by textual prompts, without the need for u…
Video Salient Object Detection via Contrastive Features and Attention Modules
Yi-Wen Chen, Xiaojie Jin, Xiaohui Shen +1
Video salient object detection aims to find the most visually distinctive objects in a video. To explore the temporal dependencies, existing methods usually resort to recurrent neu…
End-to-end Multi-modal Video Temporal Grounding
Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang
We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from…