collaborators

8 papers

cs.CV2026

LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion

Zengqun Zhao, Ziquan Liu, Yu Cao +5

The recent success of inference-time scaling in large language models has inspired similar explorations in video diffusion. In particular, motivated by the existence of "golden noi…

cs.CV2026

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

Zengqun Zhao, Yanzuo Lu, Ziquan Liu +3

Autoregressive (AR) video diffusion has recently emerged as a promising paradigm for long video generation, enabling causal synthesis beyond the limits of bidirectional models. To…

cs.CV2026

CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning

Marios Krestenitis, Christos Tzelepis, Konstantinos Ioannidis +5

Visual-Language Models (VLMs) have achieved remarkable progress in image captioning, visual question answering, and visual reasoning. Yet they remain prone to vision-language misal…

cs.LG2026

Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis

Chen Feng, Zhuo Zhi, Zhao Huang +5

Statistically consistent methods based on the noise transition matrix () offer a theoretically grounded solution to Learning with Noisy Labels (LNL), with guarantees of converge…

cs.LG2026

Data Matters Most: Auditing Social Bias in Contrastive Vision Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

Vision-language models (VLMs) deliver strong zero-shot recognition but frequently inherit social biases from their training data. We systematically disentangle three design factors…

cs.CL2025

Breaking Language Barriers or Reinforcing Bias? A Study of Gender and Racial Disparities in Multilingual Contrastive Vision Language Models

Zahraa Al Sahili, Ioannis Patras, Matthew Purver

Multilingual vision-language models (VLMs) promise universal image-text retrieval, yet their social biases remain underexplored. We perform the first systematic audit of four publi…