2 papers
cs.CV2026
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
Akhil Ramachandran, Ankit Arun, Ashish Shenoy +8
Video Large Language Models (Video LLMs) have shown remarkable progress in understanding and reasoning about visual content, particularly in tasks involving text recognition and te…
cs.CV2024
EgoQR: Efficient QR Code Reading in Egocentric Settings
Mohsen Moslehpour, Yichao Lu, Pierce Chuang +7
QR codes have become ubiquitous in daily life, enabling rapid information exchange. With the increasing adoption of smart wearable devices, there is a need for efficient, and frict…