1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.CV2025
RAVU: Retrieval Augmented Video Understanding with Compositional Reasoning over Graph
Sameer Malik, Moyuru Yamada, Ayush Singh +1
Comprehending long videos remains a significant challenge for Large Multi-modal Models (LMMs). Current LMMs struggle to process even minutes to hours videos due to their lack of ex…
cs.CV2024★ 1 cited
GLoD: Composing Global Contexts and Local Details in Image Generation
Moyuru Yamada
Diffusion models have demonstrated their capability to synthesize high-quality and diverse images from textual prompts. However, simultaneous control over both global contexts (e.g…