12 papers
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
Mingju Gao, Jingkai Zhou, Kun Gai +2
Fréchet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-level diffusion and flow-match…
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
Junhao Chen, Mingjin Chen, Jingjia Mao +12
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been meas…
Design and Performance of 220 and 270 GHz Bandpass Filters for BICEP Array
A. Steiger, The BICEP/Keck Collaboration, P. A. R. Ade +90
The BICEP Array (BA) is the latest in the BICEP/ Keck series of experiments that aim to measure the polarization of the cosmic microwave background (CMB) with small aperture polari…
PhysRAG: Enhancing Physics-Awareness in Video Generation via Retrieval-Augmented Generation
Kexu Cheng, Zicheng Liu, Mingju Gao +2
Developing physically aware video generation models remains a significant challenge due to the difficulty in capturing diverse physical phenomena, such as thermal dynamics, mechani…
SAID: Accelerating Diffusion-Based Language Models via Scaffold-Aware Iterative Decoding
Na Li, Chengda Wang, Mingju Gao +1
Diffusion large language models (DLLMs) enable non-autoregressive generation by iteratively denoising corrupted token sequences with bidirectional context. Despite their ability to…
SemBlock: Semantic Boundary Dynamic Blocks for Diffusion LLMs
Xinrui Song, Zhuoran Wang, Mingju Gao +1
Diffusion language models (DLMs) generate text through iterative denoising, and blockwise decoding improves their practicality by committing tokens in local blocks. However, existi…