4 papers
OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models
Zhaoyuan He, Muhammad Muaz, Lili Qiu
High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative quality but remain expensive…
VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
Kyoungjun Park, Yifan Yang, Juheon Yi +6
The rapid proliferation of AI-generated video necessitates robust detection tools that offer both high accuracy and human-interpretable explanations. While existing MLLM-based dete…
Joint Optimization of Handoff and Video Rate in LEO Satellite Networks
Kyoungjun Park, Zhiyuan He, Cheng Luo +4
Low Earth Orbit (LEO) satellite communication is a promising approach to providing Internet connectivity to users in many remote areas. As videos are likely to account for most tra…
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
Muhammad Muaz, Nathan Paull, Jahnavi Malagavalli
This paper presents an innovative approach to address the challenges of translating multi-modal emotion recognition models to a more practical and resource-efficient uni-modal coun…