collaborators

6 papers

cs.CV2026

Shot-Aware Frame Sampling for Video Understanding

Mengyu Zhao, Di Fu, Yongyu Xie +4

Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…

cs.CV2026

FrankenMotion: Part-level Human Motion Generation and Composition

Chuqiao Li, Xianghui Xie, Yong Cao +2

Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptio…

cs.CV2025

RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification

Meilong Xu, Di Fu, Jiaxing Zhang +7

Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particul…

cs.LG2025

MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models

Haoyu He, Katrin Renz, Yong Cao +1

Diffusion language models, as a promising alternative to traditional autoregressive (AR) models, enable faster generation and richer conditioning on bidirectional context. However,…

cs.CL2025

Scholar Inbox: Personalized Paper Recommendations for Scientists

Markus Flicke, Glenn Angrabeit, Madhav Iyengar +10

Scholar Inbox is a new open-access platform designed to address the challenges researchers face in staying current with the rapidly expanding volume of scientific literature. We pr…

cs.CL2025

Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation

Steffen Eger, Yong Cao, Jennifer D'Souza +11

With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. An emerging ecosystem of models and tools aims to su…