8 papers
NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models
Yong Cao, Chuqiao Li, Xianghui Xie +2
Reliable evaluation of human motion understanding is fundamental to advancing embodied AI, robotics, and animation. However, existing benchmarks suffer from coarse semantic granula…
UniComp: A Unified Evaluation of Large Language Model Compression via Pruning, Quantization and Distillation
Jonathan von Rad, Yong Cao, Andreas Geiger
Model compression is increasingly essential for deploying large language models (LLMs), yet existing comparative studies largely focus on pruning and quantization evaluated primari…
Shot-Aware Frame Sampling for Video Understanding
Mengyu Zhao, Di Fu, Yongyu Xie +4
Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
Steffen Eger, Yong Cao, Jennifer D'Souza +11
With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. An emerging ecosystem of models and tools aims to su…
FrankenMotion: Part-level Human Motion Generation and Composition
Chuqiao Li, Xianghui Xie, Yong Cao +2
Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptio…
RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
Meilong Xu, Di Fu, Jiaxing Zhang +7
Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particul…