32 citations · 68 across the 10 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Document Haystack: A Long Context Multimodal Image/Document Understanding Vision LLM Benchmark
Goeric Huybrechts, Srikanth Ronanki, Sai Muralidhar Jayanthi +2
The proliferation of multimodal Large Language Models has significantly advanced the ability to analyze and understand complex data inputs from different modalities. However, the p…
cs.CV2024
Adaptive Video Understanding Agent: Enhancing efficiency with dynamic frame sampling and feedback-driven reasoning
Sullam Jeoung, Goeric Huybrechts, Bhavana Ganesh +2
Understanding long-form video content presents significant challenges due to its temporal complexity and the substantial computational resources required. In this work, we propose…