5 papers
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
Qirui Li, Jinkun Hao, Yibo Li +3
Recent advances in physics-grounded video generation leverage physics simulation as a physical prior to guide video synthesis toward physically plausible outcomes. The simulation p…
Training-Free Object-Background Compositional T2I via Dynamic Spatial Guidance and Multi-Path Pruning
Yang Deng, David Mould, Paul L. Rosin +1
Existing text-to-image diffusion models, while excelling at subject synthesis, exhibit a persistent foreground bias that treats the background as a passive and under-optimized bypr…
Canonical Pose Reconstruction from Single Depth Image for 3D Non-rigid Pose Recovery on Limited Datasets
Fahd Alhamazani, Yu-Kun Lai, Paul L. Rosin
3D reconstruction from 2D inputs, especially for non-rigid objects like humans, presents unique challenges due to the significant range of possible deformations. Traditional method…
Multi-task Geometric Estimation of Depth and Surface Normal from Monocular 360° Images
Kun Huang, Fang-Lue Zhang, Fangfang Zhang +3
Geometric estimation is required for scene understanding and analysis in panoramic 360° images. Current methods usually predict a single feature, such as depth or surface normal.…
AttentionPainter: An Efficient and Adaptive Stroke Predictor for Scene Painting
Yizhe Tang, Yue Wang, Teng Hu +5
Stroke-based Rendering (SBR) aims to decompose an input image into a sequence of parameterized strokes, which can be rendered into a painting that resembles the input image. Recent…