From the 2 of 69 linked papers with an AI index.
28 papers · 1 filter
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition
Jooyeol Yun, Jintae Park, Hyesu Lim +3
Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering…
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Hoiyeong Jin, Hyojin Jang, Junha Hyung +6
Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D s…
Probability-Conserving Flow Guidance
Parsa Esmati, Junha Hyung, Amirhossein Dadashzadeh +2
Diffusion and flow-based generative models dominate visual synthesis, with guidance aligning samples to user input and improving perceptual quality. However, Classifier-Free Guidan…
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
Junha Song, Byeongho Heo, Geonmo Gu +3
When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast…
AHS: Adaptive Head Synthesis via Synthetic Data Augmentations
Taewoong Kang, Hyojin Jang, Sohyun Jeong +4
Recent digital media advancements have created increasing demands for sophisticated portrait manipulation techniques, particularly head swapping, where one's head is seamlessly int…
RL makes MLLMs see better than SFT
Junha Song, Sangdoo Yun, Dongyoon Han +2
A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its immense parameter scale and remarka…