activity
20192026
most citedEnd-to-end Multi-modal Video Temporal Grounding

21 citations · 39 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2026

CAViAR: A Causal Video Dataset for Fine-Grained Accident Reasoning in Real-World Scenarios

Sparsh Garg, Yi-Wen Chen, Vijay Kumar B G +1

While modern autonomous driving systems excel at perception tasks such as object detection and trajectory prediction, they lack the high-level causal reasoning required to interpre…

cs.CV2025

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai +2

Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when co…

cs.CV2024

High-Quality Entity Segmentation and Grounding

Lu Qi, Yi-Wen Chen, Tao Zhang +4

In this work, we propose ESG, a pipeline for high-quality entity segmentation and grounding supported by a new dataset EntitySeg. At first, the proposed dataset naming EntitySeg co…

cs.CV2023

Text-Driven Image Editing via Learnable Regions

Yuanze Lin, Yi-Wen Chen, Yi-Hsuan Tsai +2

Language has emerged as a natural interface for image editing. In this paper, we introduce a method for region-based image editing driven by textual prompts, without the need for u…

cs.CV20212 cited

Video Salient Object Detection via Contrastive Features and Attention Modules

Yi-Wen Chen, Xiaojie Jin, Xiaohui Shen +1

Video salient object detection aims to find the most visually distinctive objects in a video. To explore the temporal dependencies, existing methods usually resort to recurrent neu…

cs.CV202121 cited

End-to-end Multi-modal Video Temporal Grounding

Yi-Wen Chen, Yi-Hsuan Tsai, Ming-Hsuan Yang

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from…