papers

Publications (31)

cs.CV2020

Deeply Aligned Adaptation for Cross-domain Object Detection

Minghao Fu, Zhenshan Xie, Wen Li +1

Cross-domain object detection has recently attracted more and more attention for real-world applications, since it helps build robust detectors adapting well to new environments. I…

cs.AI2026

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations

Minghao Fu, Fan Feng, Nicklas Hansen +1

World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control. Such representations are o…

cs.LG2025

Online Time Series Forecasting with Theoretical Guarantees

Zijian Li, Changze Zhou, Minghao Fu +6

This paper is concerned with online time series forecasting, where unknown distribution shifts occur over time, i.e., latent variables influence the mapping from historical to futu…

cs.CV2023

ESTISR: Adapting Efficient Scene Text Image Super-resolution for Real-Scenes

Minghao Fu, Xin Man, Yihan Xu +1

While scene text image super-resolution (STISR) has yielded remarkable improvements in accurately recognizing scene text, prior methodologies have placed excessive emphasis on opti…

cs.CV2025

Ovis-Image Technical Report

Guo-Hua Wang, Liangfu Cao, Tianyu Cui +8

We introduce , a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational c…

cs.CV2025

CHATS: Combining Human-Aligned Optimization and Test-Time Sampling for Text-to-Image Generation

Minghao Fu, Guo-Hua Wang, Liangfu Cao +4

Diffusion models have emerged as a dominant approach for text-to-image generation. Key components such as the human preference alignment and classifier-free guidance play a crucial…