2 citations · 3 across the 9 of their papers we have counts for
3 papers · 1 filter
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
Shivam Singh, Saptarshi Majumder, Pratik Prabhanjan Brahma +2
Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We introduce pause-and-think-T, a reasoning-centr…
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
Aditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma +2
Vision-language models (VLMs) have achieved remarkable multimodal understanding and reasoning capabilities, yet remain computationally expensive due to dense visual tokenization. E…
Structured Memory based Deep Model to Detect as well as Characterize Novel Inputs
Pratik Prabhanjan Brahma, Qiuyuan Huang, Dapeng Wu
While deep learning has pushed the boundaries in various machine learning tasks, the current models are still far away from replicating many functions that a normal human brain can…