6 papers
Balancing Understanding and Generation in Discrete Diffusion Models
Yue Liu, Yuzhong Zhao, Zheyong Xie +5
In discrete generative modeling, two dominant paradigms demonstrate divergent capabilities: Masked Diffusion Language Models (MDLM) excel at semantic understanding and zero-shot ge…
Vision Calorimeter for High-Energy Particle Detection
Hongtian Yu, Yangu Li, Yunfan Liu +3
In high-energy physics, estimating anti-neutron parameters (position and momentum) using the electromagnetic calorimeter (EMC) is crucial but challenging. To conquer this challenge…
Thinking with Images via Self-Calling Agent
Wenxi Yang, Yuzhong Zhao, Fang Wan +1
Thinking-with-images paradigms have showcased remarkable visual reasoning capability by integrating visual information as dynamic elements into the Chain-of-Thought (CoT). However,…
Geometric-Mean Policy Optimization
Yuzhong Zhao, Yue Liu, Junpeng Liu +9
Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unf…
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
Mu Zhang, Yunfan Liu, Yue Liu +2
Existing image synthesis methods for natural scenes focus primarily on foreground control, often reducing the background to simplistic textures. Consequently, these approaches tend…
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
Feng Liu, Shiwei Zhang, Xiaofeng Wang +6
As a fundamental backbone for video generation, diffusion models are challenged by low inference speed due to the sequential nature of denoising. Previous methods speed up the mode…