papers

Publications (12)

cs.CV2024

Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

Kanchana Ranasinghe, Satya Narayan Shukla, Omid Poursaeed +2

Integration of Large Language Models (LLMs) into visual domain tasks, resulting in visual-LLMs (V-LLMs), has enabled exceptional performance in vision-language tasks, particularly…

cs.CV2015

Implicit Sparse Code Hashing

Tsung-Yu Lin, Tsung-Wei Ke, Tyng-Luh Liu

We address the problem of converting large-scale high-dimensional image data into binary codes so that approximate nearest-neighbor search over them can be efficiently performed. D…

cs.CV2018

Second-order Democratic Aggregation

Tsung-Yu Lin, Subhransu Maji, Piotr Koniusz

Aggregated second-order features extracted from deep convolutional networks have been shown to be effective for texture generation, fine-grained recognition, material classificatio…

cs.CV2017

Bilinear CNNs for Fine-grained Visual Recognition

Tsung-Yu Lin, Aruni RoyChowdhury, Subhransu Maji

We present a simple and effective architecture for fine-grained visual recognition called Bilinear Convolutional Neural Networks (B-CNNs). These networks represent an image as a po…

cs.LG2026

A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning

Shashank Gupta, Chaitanya Ahuja, Tsung-Yu Lin +4

Reinforcement learning (RL)-based fine-tuning has emerged as a powerful approach for aligning diffusion models with black-box objectives. Proximal policy optimization (PPO) is a po…

cs.CV2026

Xray-Visual Models: Scaling Vision models on Industry Scale Data

Shlok Mishra, Tsung-Yu Lin, Linda Wang +24

We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social media data. Our model leverages over 15 b…

cs.CV2022

Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive Learning

Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed +4

We introduce Patch Aligned Contrastive Learning (PACL), a modified compatibility function for CLIP's contrastive loss, intending to train an alignment between the patch tokens of t…

cs.CV2022

Raising the Bar on the Evaluation of Out-of-Distribution Detection

Jishnu Mukhoti, Tsung-Yu Lin, Bor-Chun Chen +4

In image classification, a lot of development has happened in detecting out-of-distribution (OoD) data. However, most OoD detection methods are evaluated on a standard set of datas…

cs.CV2017

Improved Bilinear Pooling with CNNs

Tsung-Yu Lin, Subhransu Maji

Bilinear pooling of Convolutional Neural Network (CNN) features [22, 23], and their compact variants [10], have been shown to be effective at fine-grained recognition, scene catego…

cs.CV2016

Visualizing and Understanding Deep Texture Representations

Tsung-Yu Lin, Subhransu Maji

A number of recent approaches have used deep convolutional neural networks (CNNs) to build texture representations. Nevertheless, it is still unclear how these models represent tex…

cs.CV2019

Visualizing and Describing Fine-grained Categories as Textures

Tsung-Yu Lin, Mikayla Timm, Chenyun Wu +1

We analyze how categories from recent FGVC challenges can be described by their textural content. The motivation is that subtle differences between species of birds or butterflies…

cs.CV2016

One-to-many face recognition with bilinear CNNs

Aruni RoyChowdhury, Tsung-Yu Lin, Subhransu Maji +1

The recent explosive growth in convolutional neural network (CNN) research has produced a variety of new architectures for deep learning. One intriguing new architecture is the bil…