works on

From the 2 of 69 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

28 papers · 1 filter

cs.CV2026

ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

Jooyeol Yun, Jintae Park, Hyesu Lim +3

Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering…

cs.CV2026

InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion

Hoiyeong Jin, Hyojin Jang, Junha Hyung +6

Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) remains challenging due to inadequate 4D s…

cs.CV2026

Probability-Conserving Flow Guidance

Parsa Esmati, Junha Hyung, Amirhossein Dadashzadeh +2

Diffusion and flow-based generative models dominate visual synthesis, with guidance aligning samples to user input and improving perceptual quality. However, Classifier-Free Guidan…

cs.CV2026

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models

Junha Song, Byeongho Heo, Geonmo Gu +3

When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their intended description. In contrast…

cs.CV2026

AHS: Adaptive Head Synthesis via Synthetic Data Augmentations

Taewoong Kang, Hyojin Jang, Sohyun Jeong +4

Recent digital media advancements have created increasing demands for sophisticated portrait manipulation techniques, particularly head swapping, where one's head is seamlessly int…

cs.CV2026

RL makes MLLMs see better than SFT

Junha Song, Sangdoo Yun, Dongyoon Han +2

A dominant assumption in Multimodal Language Model (MLLM) research is that its performance is largely inherited from the LLM backbone, given its immense parameter scale and remarka…