7 papers
Low-Bitrate Video Compression through Semantic-Conditioned Diffusion
Lingdong Wang, Guan-Ming Su, Divya Kothandaraman +3
Traditional video codecs optimized for pixel fidelity collapse at ultra-low bitrates and produce severe artifacts. This failure arises from a fundamental misalignment between pixel…
Zero-Shot Personalized Camera Motion Control for Image-to-Video Synthesis
Pooja Guhan, Divya Kothandaraman, Geonsun Lee +3
Specifying nuanced and compelling camera motion remains a significant hurdle for non-expert creators using generative tools, creating an "expressive gap" where generic text prompts…
SyncTrack4D: Cross-Video Motion Alignment and Video Synchronization for Multi-Video 4D Gaussian Splatting
Yonghan Lee, Tsung-Wei Huang, Shiv Gehlot +3
Modeling dynamic 3D scenes is challenging due to their high-dimensional nature, which requires aggregating information from multiple views to reconstruct time-evolving 3D geometry…
INT-DTT+: Low-Complexity Data-Dependent Transforms for Video Coding
Samuel Fernández-Menduiña, Eduardo Pavez, Antonio Ortega +4
Discrete trigonometric transforms (DTTs), such as the DCT-2 and the DST-7, are widely used in video codecs for their balance between coding performance and computational efficiency…
SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation
Guangzhi Su, Shuchang Huang, Yutong Ke +3
Multimodal large language models (MLLMs) have achieved impressive performance across diverse tasks by jointly reasoning over textual and visual inputs. Despite their success, these…
Parametric Shadow Control for Portrait Generation in Text-to-Image Diffusion Models
Haoming Cai, Tsung-Wei Huang, Shiv Gehlot +4
Text-to-image diffusion models excel at generating diverse portraits, but lack intuitive shadow control. Existing editing approaches, as post-processing, struggle to offer effectiv…