Publications (12)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed +2
Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly…
Implicit Sparse Code Hashing
Tsung-Yu Lin, Tsung-Wei Ke, Tyng-Luh Liu
We address the problem of converting large-scale high-dimensional image data into binary codes so that approximate nearest-neighbor search over them can be efficiently performed. D…
Second-order Democratic Aggregation
Tsung-Yu Lin, Subhransu Maji, Piotr Koniusz
Aggregated second-order features extracted from deep convolutional networks have been shown to be effective for texture generation, fine-grained recognition, material classificatio…
Bilinear CNNs for Fine-grained Visual Recognition
Tsung-Yu Lin, Aruni RoyChowdhury, Subhransu Maji
We present a simple and effective architecture for fine-grained visual recognition called Bilinear Convolutional Neural Networks (B-CNNs). These networks represent an image as a po…
A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning
Shashank Gupta, Chaitanya Ahuja, Tsung-Yu Lin +4
Reinforcement learning (RL)-based fine-tuning has emerged as a powerful approach for aligning diffusion models with black-box objectives. Proximal policy optimization (PPO) is a po…
Xray-Visual Models: Scaling Vision models on Industry Scale Data
Shlok Mishra, Tsung-Yu Lin, Linda Wang +24
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…
Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed +4
We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of t…
Raising the Bar on the Evaluation of Out-of-Distribution Detection
Jishnu Mukhoti, Tsung-Yu Lin, Bor-Chun Chen +4
In image classification, a lot of development has happened in detecting out-of-distribution (OoD) data. However, most OoD detection methods are evaluated on a standard set of datas…
Improved Bilinear Pooling with CNNs
Tsung-Yu Lin, Subhransu Maji
Bilinear pooling of Convolutional Neural Network (CNN) features [22, 23], and their compact variants [10], have been shown to be effective at fine-grained recognition, scene catego…
Visualizing and Understanding Deep Texture Representations
Tsung-Yu Lin, Subhransu Maji
A number of recent approaches have used deep convolutional neural networks (CNNs) to build texture representations. Nevertheless, it is still unclear how these models represent tex…
Visualizing and Describing Fine-grained Categories as Textures
Tsung-Yu Lin, Mikayla Timm, Chenyun Wu +1
We analyze how categories from recent FGVC challenges can be described by their textural content. The motivation is that subtle differences between species of birds or butterflies…
One-to-many face recognition with bilinear CNNs
Aruni RoyChowdhury, Tsung-Yu Lin, Subhransu Maji +1
The recent explosive growth in convolutional neural network (CNN) research has produced a variety of new architectures for deep learning. One intriguing new architecture is the bil…