5 papers
Real-Time Human Frontal View Synthesis from a Single Image
Fangyu Lin, Yingdong Hu, Lunjie Zhu +4
Photorealistic human novel view synthesis from a single image is crucial for democratizing immersive 3D telepresence, eliminating the need for complex multi-camera setups. However,…
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
Yufei Xue, Yushi Huang, Jiawei Shao +4
Post-training quantization (PTQ) has emerged as an effective technique for compressing large models and accelerating inference without retraining. While PTQ has been extensively st…
Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video Generation
Lunjie Zhu, Yushi Huang, Xingtong Ge +5
Latent diffusion models have enabled high-quality video synthesis, yet their inference remains costly and time-consuming. As diffusion transformers become increasingly efficient, t…
R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
Lu Zhu, Tiantian Geng, Yangye Chen +3
Recently, rapid advancements have been made in multimodal large language models (MLLMs), especially in video understanding tasks. However, current research focuses on simple video…
Conditional Chemical Language Models are Versatile Tools in Drug Discovery
Lu Zhu, Emmanuel Noutahi
Generative chemical language models (CLMs) have demonstrated strong capabilities in molecular design, yet their impact in drug discovery remains limited by the absence of reliable…