58 citations · 106 across the 28 of their papers we have counts for
44 papers
RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement
Zi-Xiang Ni, Bo-Lun Huang, Teng-Fang Hsiao +2
Text-to-image (T2I) diffusion models have achieved striking progress but still struggle to synthesize rare concepts involving unusual attribute-object pairings, often resulting in…
IDAG-Edit: Multi-Object Video Editing via Instance-Decoupled Attention and Guidance
Yuan-Zhih Lin, Huu-Thang Nguyen, Huu-Phu Do +2
Diffusion-based video editing has made significant progress; however, achieving precise and temporally consistent object-level control, especially in multi-object scenarios, remain…
HyperPatch: Sequential Knowledge Editing Under n-ary Structural Drift
Yu-Kai Chan, Wen-Sheng Lien, Dong-Ting Yao +4
Large Language Models (LLMs) rely on Knowledge Editing (KE) to maintain temporal validity, yet real-world knowledge is inherently n-ary. We demonstrate that in non-stationary envir…
HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented Generation
Wen-Sheng Lien, Yu-Kai Chan, Hao-Lung Hsiao +5
Graph-based retrieval-augmented generation (RAG) methods, typically built on knowledge graphs (KGs) with binary relational facts, have shown promise in multi-hop open-domain QA. Ho…
CookAnything: A Framework for Flexible and Consistent Multi-Step Recipe Image Generation
Ruoxuan Zhang, Bin Wen, Hongxia Xie +5
Cooking is a sequential and visually grounded activity, where each step such as chopping, mixing, or frying carries both procedural logic and visual semantics. While recent diffusi…
COACH: Collaborative Agents for Contextual Highlighting -- A Multi-Agent Framework for Sports Video Analysis
Tsz-To Wong, Ching-Chun Huang, Hong-Han Shuai
Intelligent sports video analysis demands a comprehensive understanding of temporal context, from micro-level actions to macro-level game strategies. Existing end-to-end models oft…