2 citations · 2 across the 4 of their papers we have counts for
4 papers
Seeing Straight: Document Orientation Detection for Efficient OCR
Suranjan Goswami, Abhinav Ravi, Raja Kolla +5
Despite significant advances in document understanding, determining the correct orientation of scanned or photographed documents remains a critical pre-processing step in the real…
Rethinking VLMs and LLMs for Image Classification
Avi Cooper, Keizo Kato, Chia-Hsien Shih +8
Visual Language Models (VLMs) are now increasingly being merged with Large Language Models (LLMs) to enable new capabilities, particularly in terms of improved interactivity and op…
Three approaches to facilitate DNN generalization to objects in out-of-distribution orientations and illuminations
Akira Sakai, Taro Sunagawa, Spandan Madan +7
The training data distribution is often biased towards objects in certain orientations and illumination conditions. While humans have a remarkable capability of recognizing objects…
Annotation Cost Reduction of Stream-based Active Learning by Automated Weak Labeling using a Robot Arm
Kanata Suzuki, Taro Sunagawa, Tomotake Sasaki +1
Stream-based active learning (AL) is an efficient training data collection method, and it is used to reduce human annotation cost required in machine learning. However, it is diffi…