11 papers
FlashSign: Pose-Free Guidance for Efficient Sign Language Video Generation
Liuzhou Zhang, Zeyu Zhang, Biao Wu +10
Sign language plays a crucial role in bridging communication gaps between the deaf and hard-of-hearing communities. However, existing sign language video generation models often re…
Twin Co-Adaptive Dialogue for Progressive Image Generation
Jianhui Wang, Yangfan He, Yan Zhong +12
Modern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in…
Generalized Radius and Integrated Codebook Transforms for Differentiable Vector Quantization
Haochen You, Heng Zhang, Hongyang He +2
Vector quantization (VQ) underpins modern generative and representation models by turning continuous latents into discrete tokens. Yet hard nearest-neighbor assignments are non-dif…
Readout-Side Bypass for Residual Hybrid Quantum-Classical Models
Guilin Zhang, Wulan Guo, Ziqi Tan +3
Quantum machine learning (QML) promises compact and expressive representations, but suffers from the measurement bottleneck - a narrow quantum-to-classical readout that limits perf…
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
Chak-Wing Mak, Guanyu Zhu, Boyi Zhang +16
Modern foundational Multimodal Large Language Models (MLLMs) and video world models have advanced significantly in mathematical, common-sense, and visual reasoning, but their grasp…
TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning
Hongyang He, Xinyuan Song, Yangfan He +5
We introduce TRiCo, a novel triadic game-theoretic co-training framework that rethinks the structure of semi-supervised learning by incorporating a teacher, two students, and an ad…