papers

Publications (32)

cs.CV2022

YORO -- Lightweight End to End Visual Grounding

Chih-Hui Ho, Srikar Appalaraju, Bhavan Jasani +2

We present YORO - a multi-modal transformer encoder-only architecture for the Visual Grounding (VG) task. This task involves localizing, in an image, an object referred via natural…

cs.CV2025

VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding

Ofir Abramovich, Niv Nayman, Sharon Fogel +7

In recent years, notable advancements have been made in the domain of visual document understanding, with the prevailing architecture comprising a cascade of vision and language mo…

cs.CV2020

ResNeSt: Split-Attention Networks

Hang Zhang, Chongruo Wu, Zhongyue Zhang +9

It is well known that featuremap attention and multi-path representation are important for visual recognition. In this paper, we present a modularized architecture, which applies t…

cs.CV2021

Document Visual Question Answering Challenge 2020

Minesh Mathew, Ruben Tito, Dimosthenis Karatzas +2

This paper presents results of Document Visual Question Answering Challenge organized as part of "Text and Documents in the Deep Learning Era" workshop, in CVPR 2020. The challenge…

cs.CV2022

Towards Weakly-Supervised Text Spotting using a Multi-Task Transformer

Yair Kittenplon, Inbal Lavi, Sharon Fogel +3

Text spotting end-to-end methods have recently gained attention in the literature due to the benefits of jointly optimizing the text detection and recognition components. Existing…

cs.CV2022

Searching for Apparel Products from Images in the Wild

Son Tran, Ming Du, Sampath Chanda +2

In this age of social media, people often look at what others are wearing. In particular, Instagram and Twitter influencers often provide images of themselves wearing different out…

cs.CV2023

DocFormerv2: Local Features for Document Understanding

Srikar Appalaraju, Peng Tang, Qi Dong +3

We propose DocFormerv2, a multi-modal transformer for Visual Document Understanding (VDU). The VDU domain entails understanding documents (beyond mere OCR predictions) e.g., extrac…

cs.CV2020

On Calibration of Scene-Text Recognition Models

Ron Slossberg, Oron Anschel, Amir Markovitz +6

In this work, we study the problem of word-level confidence calibration for scene-text recognition (STR). Although the topic of confidence calibration has been an active research a…

cs.CV2024

NAVERO: Unlocking Fine-Grained Semantics for Video-Language Compositionality

Chaofan Tao, Gukyeong Kwon, Varad Gunjal +7

We study the capability of Video-Language (VidL) models in understanding compositions between objects, attributes, actions and their relations. Composition understanding becomes pa…

cs.CV2025

R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding

Joonhyung Park, Peng Tang, Sagnik Das +4

Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by advances in large Vision Language…

cs.CV2024

DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models

Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6

Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…

cs.CV2018

Sampling Matters in Deep Embedding Learning

Chao-Yuan Wu, R. Manmatha, Alexander J. Smola +1

Deep embeddings answer one simple question: How similar are two images? Learning these embeddings is the bedrock of verification, zero-shot learning, and visual search. The most pr…

cs.CV2020

A Comprehensive Study of Deep Video Action Recognition

Yi Zhu, Xinyu Li, Chunhui Liu +7

Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks t…

cs.CV2023

PolyFormer: Referring Image Segmentation as Sequential Polygon Generation

Jiang Liu, Hui Ding, Zhaowei Cai +4

In this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image segmentation is formulated as sequential polygon generation, and the…

cs.CV2023

DEED: Dynamic Early Exit on Decoder for Accelerating Encoder-Decoder Transformer Models

Peng Tang, Pengkai Zhu, Tian Li +3

Encoder-decoder transformer models have achieved great success on various vision-language (VL) tasks, but they suffer from high inference latency. Typically, the decoder takes up m…

cs.CV2023

DocTr: Document Transformer for Structured Information Extraction in Documents

Haofu Liao, Aruni RoyChowdhury, Weijian Li +6

We present a new formulation for structured information extraction (SIE) from visually rich documents. It aims to address the limitations of existing IOB tagging or graph-based for…

eess.IV2020

Saliency Driven Perceptual Image Compression

Yash Patel, Srikar Appalaraju, R. Manmatha

This paper proposes a new end-to-end trainable model for lossy image compression, which includes several novel components. The method incorporates 1) an adequate perceptual similar…

cs.CV2024

On the Scalability of Diffusion-based Text-to-Image Generation

Hao Li, Yang Zou, Ying Wang +7

Scaling up model and data size has been quite successful for the evolution of LLMs. However, the scaling law for the diffusion based text-to-image (T2I) models is not fully explore…

cs.CV2021

LaTr: Layout-Aware Transformer for Scene-Text VQA

Ali Furkan Biten, Ron Litman, Yusheng Xie +2

We propose a novel multimodal architecture for Scene Text Visual Question Answering (STVQA), named Layout-Aware Transformer (LaTr). The task of STVQA requires models to reason over…

cs.AI2025

The Amazon Nova Family of Models: Technical Report and Model Card

Amazon AGI, Aaron Langford, Aayush Shah +783

We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance. Amazon Nova Pro is a highl…

cs.CV2024

Mixed-Query Transformer: A Unified Image Segmentation Architecture

Pei Wang, Zhaowei Cai, Hao Yang +3

Existing unified image segmentation models either employ a unified architecture across multiple tasks but use separate weights tailored to each dataset, or apply a single set of we…

cs.CV2023

Multiple-Question Multiple-Answer Text-VQA

Peng Tang, Srikar Appalaraju, R. Manmatha +2

We present Multiple-Question Multiple-Answer (MQMA), a novel approach to do text-VQA in encoder-decoder transformer models. The text-VQA task requires a model to answer a question…

cs.CV2022

GLASS: Global to Local Attention for Scene-Text Spotting

Roi Ronen, Shahar Tsiper, Oron Anschel +3

In recent years, the dominant paradigm for text spotting is to combine the tasks of text detection and recognition into a single end-to-end framework. Under this paradigm, both tas…

cs.CV2020

Improving Semantic Segmentation via Self-Training

Yi Zhu, Zhongyue Zhang, Chongruo Wu +6

Deep learning usually achieves the best results with complete supervision. In the case of semantic segmentation, this means that large amounts of pixelwise annotations are required…

cs.CV2024

Efficient Scaling of Diffusion Transformers for Text-to-Image Generation

Hao Li, Shamit Lal, Zhiheng Li +9

We empirically study the scaling properties of various Diffusion Transformers (DiTs) for text-to-image generation by performing extensive and rigorous ablations, including training…

cs.CV2018

Compressed Video Action Recognition

Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu +3

Training robust deep video representations has proven to be much more challenging than learning deep image representations. This is in part due to the enormous size of raw video st…

eess.IV2019

Human Perceptual Evaluations for Image Compression

Yash Patel, Srikar Appalaraju, R. Manmatha

Recently, there has been much interest in deep learning techniques to do image compression and there have been claims that several of these produce better results than engineered c…

eess.IV2019

Deep Perceptual Compression

Yash Patel, Srikar Appalaraju, R. Manmatha

Several deep learned lossy compression techniques have been proposed in the recent literature. Most of these are optimized by using either MS-SSIM (multi-scale structural similarit…

cs.CV2020

SCATTER: Selective Context Attentional Scene Text Recognizer

Ron Litman, Oron Anschel, Shahar Tsiper +3

Scene Text Recognition (STR), the task of recognizing text against complex image backgrounds, is an active area of research. Current state-of-the-art (SOTA) methods still struggle…

cs.CV2020

Sequence-to-Sequence Contrastive Learning for Text Recognition

Aviad Aberdam, Ron Litman, Shahar Tsiper +5

We propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence…

cs.CV2023

SimCon Loss with Multiple Views for Text Supervised Semantic Segmentation

Yash Patel, Yusheng Xie, Yi Zhu +2

Learning to segment images purely by relying on the image-text alignment from web data can lead to sub-optimal performance due to noise in the data. The noise comes from the sample…

cs.CV2021

DocFormer: End-to-End Transformer for Document Understanding

Srikar Appalaraju, Bhavan Jasani, Bhargava Urala Kota +2

We present DocFormer -- a multi-modal transformer based architecture for the task of Visual Document Understanding (VDU). VDU is a challenging problem which aims to understand docu…