Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
Akhil Ramachandran, Ankit Arun, Ashish Shenoy +8
Video Large Language Models (Video LLMs) have shown remarkable progress in understanding and reasoning about visual content, particularly in tasks involving text recognition and te…
cs.CV2024
EgoQR: Efficient QR Code Reading in Egocentric Settings
Mohsen Moslehpour, Yichao Lu, Pierce Chuang +7
QR codes have become ubiquitous in daily life, enabling rapid information exchange. With the increasing adoption of smart wearable devices, there is a need for efficient, and frict…
cs.CV2024
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
Ashish Shenoy, Yichao Lu, Srihari Jayakumar +11
We introduce Lumos, the first end-to-end multimodal question-answering system with text understanding capabilities. At the core of Lumos is a Scene Text Recognition (STR) component…