8 papers
Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models
Yawen Shao, Jie Xiao, Kai Zhu +6
Reinforcement learning (RL) holds immense promise for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, progress is fundamentally constraine…
Tail-Aware HiFloat4: W4A4 Post-Training Quantization for Wan2.2
Zhanfeng Feng, Shuai Guo, Xin Di +3
This report describes Tail-Aware HiFloat4, our submission to the low-bit text-to-video generation quantization challenge. Our method adapts the public ViDiT-Q post-training quantiz…
Dual-Pathway Geometry-Aware MLLM for Spatial Intelligence
Yufei Zheng, Xuhan Zhu, Zide Liu +9
Spatial understanding of the physical world from 2D visual inputs hinges on two complementary forms of geometric knowledge: holistic 3D structural perception and fine-grained metri…
TIE: Time Interval Encoding for Video Generation over Events
Zhilei Shu, Shangwen Zhu, Zihang Liang +10
Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and ove…
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Chenfeng Wang, Wei He, Xuhan Zhu +10
In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…
Gloria: Consistent Character Video Generation via Content Anchors
Yuhang Yang, Fan Zhang, Huaijin Pi +5
Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Ex…