Publications (23)
Sparse Fusion for Multimodal Transformers
Yi Ding, Alex Rich, Mason Wang +4
Multimodal classification is a core task in human-centric machine learning. We observe that information is highly complementary across modalities, thus unimodal information can be…
Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning
Sherry X. Chen, Misha Sra, Pradeep Sen
Although natural language instructions offer an intuitive way to guide automated image editing, deep-learning models often struggle to achieve high-quality results, largely due to…
Noise-Aware Video Saliency Prediction
Ekta Prashnani, Orazio Gallo, Joohwan Kim +3
We tackle the problem of predicting saliency maps for videos of dynamic scenes. We note that the accuracy of the maps reconstructed from the gaze data of a fixed number of observer…
Patch-Based Image Hallucination for Super Resolution with Detail Reconstruction from Similar Sample Images
Chieh-Chi Kao, Yuxiang Wang, Jonathan Waltman +1
Image hallucination and super-resolution have been studied for decades, and many approaches have been proposed to upsample low-resolution images using information from the images t…
SiCo: An Interactive Size-Controllable Virtual Try-On Approach for Informed Decision-Making
Sherry X. Chen, Alex Christopher Lim, Yimeng Liu +2
Virtual try-on (VTO) applications aim to replicate the in-store shopping experience and enhance online shopping by enabling users to interact with garments. However, many existing…
Cosmic-CoNN: A Cosmic Ray Detection Deep-Learning Framework, Dataset, and Toolkit
Chengyuan Xu, Curtis McCully, Boning Dong +2
Rejecting cosmic rays (CRs) is essential for the scientific interpretation of CCD-captured data, but detecting CRs in single-exposure images has remained challenging. Conventional…
VoRTX: Volumetric 3D Reconstruction With Transformers for Voxelwise View Selection and Fusion
Noah Stier, Alexander Rich, Pradeep Sen +1
Recent volumetric 3D reconstruction methods can produce very accurate results, with plausible geometry even for unobserved surfaces. However, they face an undesirable trade-off whe…
Improving the Resolution of CNN Feature Maps Efficiently with Multisampling
Shayan Sadigh, Pradeep Sen
We describe a new class of subsampling techniques for CNNs, termed multisampling, that significantly increases the amount of information kept by feature maps through subsampling la…
AID-AppEAL: Automatic Image Dataset and Algorithm for Content Appeal Enhancement and Assessment Labeling
Sherry X. Chen, Yaron Vaxman, Elad Ben Baruch +4
We propose Image Content Appeal Assessment (ICAA), a novel metric that quantifies the level of positive interest an image's content generates for viewers, such as the appeal of foo…
TiNO-Edit: Timestep and Noise Optimization for Robust Diffusion-Based Image Editing
Sherry X. Chen, Yaron Vaxman, Elad Ben Baruch +5
Despite many attempts to leverage pre-trained text-to-image models (T2I) like Stable Diffusion (SD) for controllable image editing, producing good predictable results remains a cha…
ANSAC: Adaptive Non-minimal Sample and Consensus
Victor Fragoso, Chris Sweeney, Pradeep Sen +1
While RANSAC-based methods are robust to incorrect image correspondences (outliers), their hypothesis generators are not robust to correct image correspondences (inliers) with posi…
3DVNet: Multi-View Depth Prediction and Volumetric Refinement
Alexander Rich, Noah Stier, Pradeep Sen +1
We present 3DVNet, a novel multi-view stereo (MVS) depth-prediction method that combines the advantages of previous depth-based and volumetric MVS approaches. Our key idea is the u…
Localization-Aware Active Learning for Object Detection
Chieh-Chi Kao, Teng-Yok Lee, Pradeep Sen +1
Active learning - a class of algorithms that iteratively searches for the most informative samples to include in a training dataset - has been shown to be effective at annotating d…
Meshlet Priors for 3D Mesh Reconstruction
Abhishek Badki, Orazio Gallo, Jan Kautz +1
Estimating a mesh from an unordered set of sparse, noisy 3D points is a challenging problem that requires carefully selected priors. Existing hand-crafted priors, such as smoothnes…
Interactive Segmentation and Visualization for Tiny Objects in Multi-megapixel Images
Chengyuan Xu, Boning Dong, Noah Stier +4
We introduce an interactive image segmentation and visualization framework for identifying, inspecting, and editing tiny objects (just a few pixels wide) in large multi-megapixel h…
PieAPP: Perceptual Image-Error Assessment through Pairwise Preference
Ekta Prashnani, Hong Cai, Yasamin Mostofi +1
The ability to estimate the perceptual error between images is an important problem in computer vision with many applications. Although it has been studied extensively, however, no…
Deep Appearance Prefiltering
Steve Bako, Pradeep Sen, Anton Kaplanyan
Physically based rendering of complex scenes can be prohibitively costly with a potentially unbounded and uneven distribution of complexity across the rendered image. The goal of a…
Prism: Semi-Supervised Multi-View Stereo with Monocular Structure Priors
Alex Rich, Noah Stier, Pradeep Sen +1
The promise of unsupervised multi-view-stereo (MVS) is to leverage large unlabeled datasets, yet current methods underperform when training on difficult data, such as handheld smar…
Bi3D: Stereo Depth Estimation via Binary Classifications
Abhishek Badki, Alejandro Troccoli, Kihwan Kim +3
Stereo-based depth estimation is a cornerstone of computer vision, with state-of-the-art methods delivering accurate results in real time. For several applications such as autonomo…
GraphMatch: Efficient Large-Scale Graph Construction for Structure from Motion
Qiaodong Cui, Victor Fragoso, Chris Sweeney +1
We present GraphMatch, an approximate yet efficient method for building the matching graph for large-scale structure-from-motion (SfM) pipelines. Unlike modern SfM pipelines that u…
Binary TTC: A Temporal Geofence for Autonomous Navigation
Abhishek Badki, Orazio Gallo, Jan Kautz +1
Time-to-contact (TTC), the time for an object to collide with the observer's plane, is a powerful tool for path planning: it is potentially more informative than the depth, velocit…
Efficient Scene Appearance Aggregation for Level-of-Detail Rendering
Yang Zhou, Tao Huang, Ravi Ramamoorthi +2
Creating an appearance-preserving level-of-detail (LoD) representation for arbitrary 3D scenes is a challenging problem. The appearance of a scene is an intricate combination of bo…
On the Relationship Between Dual Photography and Classical Ghost Imaging
Pradeep Sen
Classical ghost imaging has received considerable attention in recent years because of its remarkable ability to image a scene without direct observation by a light-detecting imagi…