From the 1 of 12 linked papers with an AI index.
12 papers
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
Jiaang Li, Chengzu Li, Zhaochong An +4
The paper investigates why multimodal large language models often ignore visual evidence, using image reconstruction and a new benchmark (WhatIfVis) to measure how well models bala…
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Mingqiao Ye, Zhaochong An, Zhitong Gao +11
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-language models, and increasingly…
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
Zhaochong An, Orest Kupyn, Théo Uscidda +5
Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting th…
Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks
Feng Qiao, Zhaochong An, Zhexiao Xiong +2
Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the orig…
Stitched Value Model for Diffusion Alignment
Hyojun Go, Hyungjin Chung, Prune Truong +8
For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challen…
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
Dan Wang, Haiyan Sun, Shan Du +4
Image super-resolution (SR) aims to reconstruct high resolution images with both high perceptual quality and low distortion, but is fundamentally limited by the perception-distorti…