collaborators

5 papers

cs.CV2026

Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration

Zihao He, Yunfeng Wu, Xinchao Wang +1

All-in-one image restoration seeks a single model that can recover images degraded by diverse and spatially non-uniform corruptions. However, many unified Transformers rely on fixe…

cs.CV2026

Flow-Map Distillation on Relation Manifolds for Image Restoration

Zihao He, Songhua Liu

Knowledge distillation for image restoration typically aligns intermediate features or relation matrices between teacher and student networks as static targets, ignoring the dynami…

cs.CL2026

Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge

Junjie Wu, Xuan Kan, Zihao He +3

Multimodal Large Language Models (MLLMs) have been widely adopted as MLLM-as-a-Judges due to their strong alignment with human judgment across various visual tasks. However, most e…

cs.CV2026

ViBe: Ultra-High-Resolution Video Synthesis Born from Pure Images

Yunfeng Wu, Hongying Cheng, Zihao He +1

Transformer-based video diffusion models rely on 3D attention over spatial and temporal tokens, which incurs quadratic time and memory complexity and makes end-to-end training for…

cs.CV2025

Towards Pixel-Level Prediction for Gaze Following: Benchmark and Approach

Feiyang Liu, Dan Guo, Jingyuan Xu +4

Following the gaze of other people and analyzing the target they are looking at can help us understand what they are thinking, and doing, and predict the actions that may follow. E…