collaborators

5 papers

cs.LG2026

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models

Haiyu Wang, Yutong Wang, Leshu Li +2

Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource…

cs.CV2025

VDOT: Efficient Unified Video Creation via Optimal Transport Distillation

Yutong Wang, Haiyu Zhang, Tianfan Xue +4

The rapid development of generative models has significantly advanced image and video applications. Among these, video creation, aimed at generating videos under various conditions…

cs.LG2025

QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models

Yutong Wang, Haiyu Wang, Sai Qian Zhang

Vision-Language Models (VLMs) are integral to tasks such as image captioning and visual question answering, but their high computational cost, driven by large memory footprints and…

cs.CL2025

WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines

Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan +48

Vision Language Models (VLMs) often struggle with culture-specific knowledge, particularly in languages other than English and in underrepresented cultural contexts. To evaluate th…

cs.CV2025

Make-A-Character 2: Animatable 3D Character Generation From a Single Image

Lin Liu, Yutong Wang, Jiahao Chen +5

This report introduces Make-A-Character 2, an advanced system for generating high-quality 3D characters from single portrait photographs, ideal for game development and digital hum…