2 citations · 2 across the 2 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
Darshan Singh, Arsha Nagrani, Kawshik Manikantan +6
Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data…
cs.CV2024
GraPE: A Generate-Plan-Edit Framework for Compositional T2I Synthesis
Ashish Goswami, Satyam Kumar Modi, Santhosh Rishi Deshineni +3
Text-to-image (T2I) generation has seen significant progress with diffusion models, enabling generation of photo-realistic images from text prompts. Despite this progress, existing…