activity
20172023
most citedFrom Satellite Imagery to Disaster Insights

50 citations · 70 across the 8 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2023★ 2 cited

DISGO: Automatic End-to-End Evaluation for Scene Text OCR

Mei-Yuh Hwang, Yangyang Shi, Ankit Ramchandani +6

This paper discusses the challenges of optical character recognition (OCR) on natural scenes, which is harder than OCR on documents due to the wild content and various image backgr…

cs.CV2023★ 2 cited

Text-Conditional Contextualized Avatars For Zero-Shot Personalization

Samaneh Azadi, Thomas Hayes, Akbar Shah +3

Recent large-scale text-to-image generation models have made significant improvements in the quality, realism, and diversity of the synthesized images and enable users to control t…

cs.CV2022

MUGEN: A Playground for Video-Audio-Text Multimodal Understanding and GENeration

Thomas Hayes, Songyang Zhang, Xi Yin +6

Multimodal video-audio-text understanding and generation can benefit from datasets that are narrow but rich. The narrowness allows bite-sized challenges that the research community…

cs.CV2021★ 3 cited

TextStyleBrush: Transfer of Text Aesthetics from a Single Example

Praveen Krishnan, Rama Kovvuri, Guan Pang +2

We present a novel approach for disentangling the content of a text image from all aspects of its appearance. The appearance representation we derive can then be applied to new con…

cs.CV2021

TextOCR: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text

Amanpreet Singh, Guan Pang, Mandy Toh +3

A crucial component for the scene text based reasoning required for TextVQA and TextCaps datasets involve detecting and recognizing text present in the images using an optical char…

cs.CV2021

A Multiplexed Network for End-to-End, Multilingual OCR

Jing Huang, Guan Pang, Rama Kovvuri +5

Recent advances in OCR have shown that an end-to-end (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods…