3 papers
cs.CV2026
LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding
Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase +3
Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Train…
cs.DC2026
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase +4
Multimodal deep learning models enable joint learning across heterogeneous data sources, including text, images, and video, but their rapid scaling introduces significant memory an…
cs.CV2025
From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
Le Zhuo, Liangbing Zhao, Sayak Paul +6
Recent text-to-image diffusion models achieve impressive visual quality through extensive scaling of training data and model parameters, yet they often struggle with complex scenes…