activity
20232026
collaborators

7 papers

cs.CV2026

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase +3

Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Train…

cs.DC2026

Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips

Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase +4

Multimodal deep learning models enable joint learning across heterogeneous data sources, including text, images, and video, but their rapid scaling introduces significant memory an…

cs.CV2025

3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes

Mahmoud Ahmed, Xiang Li, Arpit Prajapati +1

Understanding objects in 3D at the part level is essential for humans and robots to navigate and interact with the environment. Current datasets for part-level 3D object understand…

cs.CV2024

InfiniBench: A Benchmark for Large Multi-Modal Models in Long-Form Movies and TV Shows

Kirolos Ataallah, Eslam Abdelrahman, Mahmoud Ahmed +4

Understanding long-form videos, such as movies and TV episodes ranging from tens of minutes to two hours, remains a significant challenge for multi-modal models. Existing benchmark…

cs.CV2024

Kestrel: 3D Multimodal LLM for Part-Aware Grounded Description

Mahmoud Ahmed, Junjie Fei, Jian Ding +2

In this paper, we introduce Part-Aware Point Grounded Description (PaPGD), a challenging task aimed at advancing 3D multimodal learning for fine-grained, part-aware segmentation gr…

cs.CV2023

3DCoMPaT: An improved Large-scale 3D Vision Dataset for Compositional Recognition

Habib Slim, Xiang Li, Yuchen Li +8

In this work, we present 3DCoMPaT, a multimodal 2D/3D dataset with 160 million rendered views of more than 10 million stylized 3D shapes carefully annotated at the part-inst…