1 citations · 1 across the 29 of their papers we have counts for
27 papers · 1 filter
Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics
Ruibo Ming, Lei Sun, Deheng Zhang +8
Many emerging video language modeling tasks require systems to move beyond clip-level abstraction and model visual content as it unfolds over extended time horizons. However, most…
Geometry-Aware Test-Time Learning for Quantitative Spatial Reasoning
Gege Zhang, Shuaicheng Niu, Gang Dai +2
Quantitative spatial reasoning in visual-language models (VLMs) aims to infer spatial distances and directional relationships among objects in 3D space from a 2D image and a natura…
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…
Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment
Geng Li, Haiwen Li, Rui Chen +3
Video aesthetic assessment (VAA) aims to predict how aesthetically pleasing a video is, yet remains far less explored than other visual assessment tasks. Its progress is hindered n…
DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen +20
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
Meng Yu, Lei Sun, Jianhao Zeng +2
Diffusion Probabilistic Models have demonstrated remarkable performance across a wide range of generative tasks. However, we have observed that these models often suffer from a Sig…