2 papers
cs.LG2026
ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison
Tianle Li, Xuyang Shen, Yan Ma +7
Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the level of individual visual cl…
cs.CV2026
Towards Scalable Pre-training of Visual Tokenizers for Generation
Jingfeng Yao, Yuda Song, Yucong Zhou +1
The quality of the latent space in visual tokenizers (e.g., VAEs) is crucial for modern generative models. However, the standard reconstruction-based training paradigm produces a l…