6 papers
Surprise2Refine: Axis-Centered Exploration-To-Refinement for Agent-Assisted Creative Scaffolding
Yuzhe You, Gromit Yeuk-Yin Chan, Shunan Guo +4
Designers require different design spaces across creative stages: broad during exploration, and targeted during refinement. Yet existing agent-driven tools assume a fixed or contin…
Dynamic Neural Graph Encoding of Inference Processes in Deep Weight Space
Di Wu, Huan Liu, Zhixiang Chi +3
The rapid advancements in using neural networks as implicit data representations have attracted significant interest in developing machine learning methods that analyze and process…
Real-time Appearance-based Gaze Estimation for Open Domains
Zhenhao Li, Zheng Liu, Seunghyun Lee +2
Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap where existing AGE models often f…
Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
Houston H. Zhang, Tao Zhang, Baoze Lin +10
User interface to code (UI2Code) aims to generate executable code that can faithfully reconstruct a given input UI. Prior work focuses largely on web pages and mobile screens, leav…
CamDirector: Towards Long-Term Coherent Video Trajectory Editing
Zhihao Shi, Kejia Yin, Weilin Wan +5
Video (camera) trajectory editing aims to synthesize new videos that follow user-defined camera paths while preserving scene content and plausibly inpainting previously unseen regi…
Few-Shot-Based Modular Image-to-Video Adapter for Diffusion Models
Zhenhao Li, Shaohan Yi, Zheng Liu +7
Diffusion models (DMs) have recently achieved impressive photorealism in image and video generation. However, their application to image animation remains limited, even when traine…