7 papers
Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
Xinxin Liu, Ming Li, Zonglin Lyu +2
Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets provide only single, holisti…
Geo: Geometry-Guided Cross-view Geo-Localization and Image Synthesis
Yancheng Zhang, Xiaohan Zhang, Guangyu Sun +3
Cross-view geo-spatial learning consists of two important tasks: Cross-View Geo-Localization (CVGL) and Cross-View Image Synthesis (CVIS), both of which rely on establishing geomet…
CPO: Condition Preference Optimization for Controllable Image Generation
Zonglin Lyu, Ming Li, Xinxin Liu +1
To enhance controllability in text-to-image generation, ControlNet introduces image-based control signals, while ControlNet++ improves pixel-level cycle consistency between generat…
TLB-VFI: Temporal-Aware Latent Brownian Bridge Diffusion for Video Frame Interpolation
Zonglin Lyu, Chen Chen
Video Frame Interpolation (VFI) aims to predict the intermediate frame (we use n to denote time in videos to avoid notation overload with the timestep in diffusion models…
Frame Interpolation with Consecutive Brownian Bridge Diffusion
Zonglin Lyu, Ming Li, Jianbo Jiao +1
Recent work in Video Frame Interpolation (VFI) tries to formulate VFI as a diffusion-based conditional image generation problem, synthesizing the intermediate frame given a random…
Tell Me Where You Are: Multimodal LLMs Meet Place Recognition
Zonglin Lyu, Juexiao Zhang, Mingxuan Lu +2
Large language models (LLMs) exhibit a variety of promising capabilities in robotics, including long-horizon planning and commonsense reasoning. However, their performance in place…