6 papers
Shot-Aware Frame Sampling for Video Understanding
Mengyu Zhao, Di Fu, Yongyu Xie +4
Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…
FrankenMotion: Part-level Human Motion Generation and Composition
Chuqiao Li, Xianghui Xie, Yong Cao +2
Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptio…
RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
Meilong Xu, Di Fu, Jiaxing Zhang +7
Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particul…
MDPO: Overcoming the Training-Inference Divide of Masked Diffusion Language Models
Haoyu He, Katrin Renz, Yong Cao +1
Diffusion language models, as a promising alternative to traditional autoregressive (AR) models, enable faster generation and richer conditioning on bidirectional context. However,…
Scholar Inbox: Personalized Paper Recommendations for Scientists
Markus Flicke, Glenn Angrabeit, Madhav Iyengar +10
Scholar Inbox is a new open-access platform designed to address the challenges researchers face in staying current with the rapidly expanding volume of scientific literature. We pr…
Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation
Steffen Eger, Yong Cao, Jennifer D'Souza +11
With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. An emerging ecosystem of models and tools aims to su…