9 papers
Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
Zheng Huang, Enpei Zhang, Weikang Qiu +7
Understanding how the brain encodes visual information is a central challenge in neuroscience and machine learning. A promising approach is to reconstruct visual stimuli, essential…
UniFixer: A Universal Reference-Guided Fixer for Diffusion-Based View Synthesis
Sihan Chen, Xiang Zhang, Yang Zhang +2
With the recent surge of generative models, diffusion-based approaches have become mainstream for view synthesis tasks, either in an explicit depth-warp-inpaint or in an implicit e…
EquiContact: A Hierarchical SE(3) Vision-to-Force Equivariant Policy for Spatially Generalizable Contact-rich Tasks
Joohwan Seo, Arvind Kruthiventy, Soomi Lee +5
This paper presents a framework for learning vision-based robotic policies for contact-rich manipulation tasks that generalize spatially across task configurations. We focus on ach…
Revisiting Lightweight Low-Light Image Enhancement: From a YUV Color Space Perspective
Hailong Yan, Shice Liu, Xiangtao Zhang +4
In the current era of mobile internet, Lightweight Low-Light Image Enhancement (L3IE) is critical for mobile devices, which faces a persistent trade-off between visual quality and…
Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel Views
Xiang Zhang, Yang Zhang, Lukas Mehl +2
Soft boundaries, like thin hairs, are commonly observed in natural and computer-generated imagery, but they remain challenging for 3D vision due to the ambiguous mixing of foregrou…
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
Yu Yang, Zhilu Zhang, Xiang Zhang +3
Interactive world models that simulate object dynamics are crucial for robotics, VR, and AR. However, it remains a significant challenge to learn physics-consistent dynamics models…