works on

From the 2 of 11 linked papers with an AI index.

collaborators

11 papers

cs.CV2026

Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion

Henglin Liu, Fangyuan Kong, Jing Wang +7

The paper introduces concentrated Implicit Preference Optimization (cIPO), a post‑training method for text‑to‑video diffusion models that derives preference signals from reconstruc…

cs.CV2026

MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control

Kaiqi Liu, Yunyao Mao, Ziqi Cai +8

MAVIN is a framework for generating multi-shot audio‑visual content with fine‑grained narrative control, using boundary‑aware attention to align temporal segments and ID‑aware prop…

cs.LG2026

Balancing Image Compression and Generation with Bootstrapped Tokenization

Haozhe Chi, Jinghan Li, Hao Jiang +4

Despite progress in image tokenization, standard methods encode redundant information by mixing all granularities within each token, thus redundancy persists between tokens. The mi…

cs.LG2026

LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching

Yao Lai, Xuyuan Xiong, Zeyue Xue +7

In semiconductor manufacturing, lithography projects circuit layouts onto silicon wafers through an optical mask. As circuit features shrink below the wavelength of light, optical…

cs.IR2026

Efficient Generative Retrieval for E-commerce Search with Semantic Cluster IDs and Expert-Guided RL

Jianbo Zhu, Xing Fang, Jing Wang +5

Generative retrieval offers a promising alternative by unifying the fragmented multi-stage retrieval process into a single end-to-end model. However, its practical adoption in indu…

cs.CV2026

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer

Qi Cai, Jingwen Chen, Chengmin Gao +22

The evolution of visual generative models has long been constrained by fragmented architectures relying on disjoint text encoders and external VAEs. In this report, we present HiDr…