collaborators

13 papers

cs.CV2026

From Pixels to Places: A Systematic Benchmark for Evaluating Image Geolocalization Ability in Large Language Models

Lingyao Li, Runlong Yu, Qikai Hu +4

Image geolocalization, the task of identifying the geographic location depicted in an image, is important for applications in crisis response, digital forensics, and location-based…

cs.CV2026

OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models

Ling Lin, Yang Bai, Heng Su +5

Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and i…

cs.LG2026

Unrewarded Exploration in Large Language Models Reveals Latent Learning from Psychology

Jian Xiong, Jingbo Zhou, Zihan Zhou +6

Latent learning, classically theorized by Tolman, shows that biological agents (e.g., rats) can acquire internal representations of their environment without rewards, enabling rapi…

cs.CV2025

Alchemist: Unlocking Efficiency in Text-to-Image Model Training via Meta-Gradient Data Selection

Kaixin Ding, Yang Zhou, Xi Chen +5

Recent advances in Text-to-Image (T2I) generative models, such as Imagen, Stable Diffusion, and FLUX, have led to remarkable improvements in visual quality. However, their performa…

cs.CV2025

RELIC: Interactive Video World Model with Long-Horizon Memory

Yicong Hong, Yiqun Mei, Chongjian Ge +11

A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approa…

cs.LG2025

Demystify Protein Generation with Hierarchical Conditional Diffusion Models

Zinan Ling, Yi Shi, Brett McKinney +3

Generating novel and functional protein sequences is critical to a wide range of applications in biology. Recent advancements in conditional diffusion models have shown impressive…