papers

Publications (23)

cs.CV2021

Sparse Fusion for Multimodal Transformers

Yi Ding, Alex Rich, Mason Wang +4

Multimodal classification is a core task in human-centric machine learning. We observe that information is highly complementary across modalities, thus unimodal information can be…

cs.CV2025

Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning

Sherry X. Chen, Misha Sra, Pradeep Sen

Although natural language instructions offer an intuitive way to guide automated image editing, deep-learning models often struggle to achieve high-quality results, largely due to…

cs.CV2021

Noise-Aware Video Saliency Prediction

Ekta Prashnani, Orazio Gallo, Joohwan Kim +3

We tackle the problem of predicting saliency maps for videos of dynamic scenes. We note that the accuracy of the maps reconstructed from the gaze data of a fixed number of observer…

cs.CV2018

Patch-Based Image Hallucination for Super Resolution with Detail Reconstruction from Similar Sample Images

Chieh-Chi Kao, Yuxiang Wang, Jonathan Waltman +1

Image hallucination and super-resolution have been studied for decades, and many approaches have been proposed to upsample low-resolution images using information from the images t…

cs.HC2025

SiCo: An Interactive Size-Controllable Virtual Try-On Approach for Informed Decision-Making

Sherry X. Chen, Alex Christopher Lim, Yimeng Liu +2

Virtual try-on (VTO) applications aim to replicate the in-store shopping experience and enhance online shopping by enabling users to interact with garments. However, many existing…

astro-ph.IM2022

Cosmic-CoNN: A Cosmic Ray Detection Deep-Learning Framework, Dataset, and Toolkit

Chengyuan Xu, Curtis McCully, Boning Dong +2

Rejecting cosmic rays (CRs) is essential for the scientific interpretation of CCD-captured data, but detecting CRs in single-exposure images has remained challenging. Conventional…

cs.CV2021

VoRTX: Volumetric 3D Reconstruction With Transformers for Voxelwise View Selection and Fusion

Noah Stier, Alexander Rich, Pradeep Sen +1

Recent volumetric 3D reconstruction methods can produce very accurate results, with plausible geometry even for unobserved surfaces. However, they face an undesirable trade-off whe…

cs.CV2023

Improving the Resolution of CNN Feature Maps Efficiently with Multisampling

Shayan Sadigh, Pradeep Sen

We describe a new class of subsampling techniques for CNNs, termed multisampling, that significantly increases the amount of information kept by feature maps through subsampling la…

cs.CV2024

AID-AppEAL: Automatic Image Dataset and Algorithm for Content Appeal Enhancement and Assessment Labeling

Sherry X. Chen, Yaron Vaxman, Elad Ben Baruch +4

We propose Image Content Appeal Assessment (ICAA), a novel metric that quantifies the level of positive interest an image's content generates for viewers, such as the appeal of foo…

cs.CV2024

TiNO-Edit: Timestep and Noise Optimization for Robust Diffusion-Based Image Editing

Sherry X. Chen, Yaron Vaxman, Elad Ben Baruch +5

Despite many attempts to leverage pre-trained text-to-image models (T2I) like Stable Diffusion (SD) for controllable image editing, producing good predictable results remains a cha…

cs.CV2017

ANSAC: Adaptive Non-minimal Sample and Consensus

Victor Fragoso, Chris Sweeney, Pradeep Sen +1

While RANSAC-based methods are robust to incorrect image correspondences (outliers), their hypothesis generators are not robust to correct image correspondences (inliers) with posi…

cs.CV2021

3DVNet: Multi-View Depth Prediction and Volumetric Refinement

Alexander Rich, Noah Stier, Pradeep Sen +1

We present 3DVNet, a novel multi-view stereo (MVS) depth-prediction method that combines the advantages of previous depth-based and volumetric MVS approaches. Our key idea is the u…

cs.CV2018

Localization-Aware Active Learning for Object Detection

Chieh-Chi Kao, Teng-Yok Lee, Pradeep Sen +1

Active learning - a class of algorithms that iteratively searches for the most informative samples to include in a training dataset - has been shown to be effective at annotating d…

cs.CV2020

Meshlet Priors for 3D Mesh Reconstruction

Abhishek Badki, Orazio Gallo, Jan Kautz +1

Estimating a mesh from an unordered set of sparse, noisy 3D points is a challenging problem that requires carefully selected priors. Existing hand-crafted priors, such as smoothnes…

cs.CV2022

Interactive Segmentation and Visualization for Tiny Objects in Multi-megapixel Images

Chengyuan Xu, Boning Dong, Noah Stier +4

We introduce an interactive image segmentation and visualization framework for identifying, inspecting, and editing tiny objects (just a few pixels wide) in large multi-megapixel h…

cs.CV2018

PieAPP: Perceptual Image-Error Assessment through Pairwise Preference

Ekta Prashnani, Hong Cai, Yasamin Mostofi +1

The ability to estimate the perceptual error between images is an important problem in computer vision with many applications. Although it has been studied extensively, however, no…

cs.GR2022

Deep Appearance Prefiltering

Steve Bako, Pradeep Sen, Anton Kaplanyan

Physically based rendering of complex scenes can be prohibitively costly with a potentially unbounded and uneven distribution of complexity across the rendered image. The goal of a…

cs.CV2024

Prism: Semi-Supervised Multi-View Stereo with Monocular Structure Priors

Alex Rich, Noah Stier, Pradeep Sen +1

The promise of unsupervised multi-view-stereo (MVS) is to leverage large unlabeled datasets, yet current methods underperform when training on difficult data, such as handheld smar…

cs.CV2020

Bi3D: Stereo Depth Estimation via Binary Classifications

Abhishek Badki, Alejandro Troccoli, Kihwan Kim +3

Stereo-based depth estimation is a cornerstone of computer vision, with state-of-the-art methods delivering accurate results in real time. For several applications such as autonomo…

cs.CV2017

GraphMatch: Efficient Large-Scale Graph Construction for Structure from Motion

Qiaodong Cui, Victor Fragoso, Chris Sweeney +1

We present GraphMatch, an approximate yet efficient method for building the matching graph for large-scale structure-from-motion (SfM) pipelines. Unlike modern SfM pipelines that u…

cs.CV2021

Binary TTC: A Temporal Geofence for Autonomous Navigation

Abhishek Badki, Orazio Gallo, Jan Kautz +1

Time-to-contact (TTC), the time for an object to collide with the observer's plane, is a powerful tool for path planning: it is potentially more informative than the depth, velocit…

cs.GR2024

Efficient Scene Appearance Aggregation for Level-of-Detail Rendering

Yang Zhou, Tao Huang, Ravi Ramamoorthi +2

Creating an appearance-preserving level-of-detail (LoD) representation for arbitrary 3D scenes is a challenging problem. The appearance of a scene is an intricate combination of bo…

physics.optics2013

On the Relationship Between Dual Photography and Classical Ghost Imaging

Pradeep Sen

Classical ghost imaging has received considerable attention in recent years because of its remarkable ability to image a scene without direct observation by a light-detecting imagi…