6 papers
IDAG-Edit: Multi-Object Video Editing via Instance-Decoupled Attention and Guidance
Yuan-Zhih Lin, Huu-Thang Nguyen, Huu-Phu Do +2
Diffusion-based video editing has made significant progress; however, achieving precise and temporally consistent object-level control, especially in multi-object scenarios, remain…
COACH: Collaborative Agents for Contextual Highlighting -- A Multi-Agent Framework for Sports Video Analysis
Tsz-To Wong, Ching-Chun Huang, Hong-Han Shuai
Intelligent sports video analysis demands a comprehensive understanding of temporal context, from micro-level actions to macro-level game strategies. Existing end-to-end models oft…
DNA: Dual-branch Network with Adaptation for Open-Set Online Handwriting Generation
Tsai-Ling Huang, Nhat-Tuong Do-Tran, Ngoc-Hoang-Lam Le +2
Online handwriting generation (OHG) enhances handwriting recognition models by synthesizing diverse, human-like samples. However, existing OHG methods struggle to generate unseen c…
DetailSemNet: Elevating Signature Verification through Detail-Semantic Integration
Meng-Cheng Shih, Tsai-Ling Huang, Yu-Heng Shih +4
Offline signature verification (OSV) is a frequently utilized technology in forensics. This paper proposes a new model, DetailSemNet, for OSV. Unlike previous methods that rely on…
Arbitrary-Resolution and Arbitrary-Scale Face Super-Resolution with Implicit Representation Networks
Yi Ting Tsai, Yu Wei Chen, Hong-Han Shuai +1
Face super-resolution (FSR) is a critical technique for enhancing low-resolution facial images and has significant implications for face-related tasks. However, existing FSR method…
Swapped Logit Distillation via Bi-level Teacher Alignment
Stephen Ekaputra Limantoro, Jhe-Hao Lin, Chih-Yu Wang +4
Knowledge distillation (KD) compresses the network capacity by transferring knowledge from a large (teacher) network to a smaller one (student). It has been mainstream that the tea…