Publications (152)
Embedding Physical Reasoning into Diffusion-Based Shadow Generation
Shilin Hu, Jingyi Xu, Akshat Dave +2
Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination. However, most existing methods operate purely in image space, leaving th…
Talking Head Generation via AU-Guided Landmark Prediction
Shao-Yu Chang, Jingyi Xu, Hieu Le +1
We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotio…
Personalized Image Descriptions from Attention Sequences
Ruoyu Xue, Hieu Le, Jingyi Xu +5
People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to s…
Predicting Visual Attention in Graphic Design Documents
Souradeep Chakraborty, Zijun Wei, Conor Kelton +4
We present a model for predicting visual attention during the free viewing of graphic design documents. While existing works on this topic have aimed at predicting static saliency…
SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel Histopathology
Saarthak Kapse, Pushpak Pati, Srijan Das +7
Introducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging, given the complexity of gigapixel slid…
Generating metamers of human scene understanding
Ritik Raina, Abe Leite, Alexandros Graikos +3
Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent underst…
CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
Prantik Howlader, Hoang Nguyen-Canh, Srijan Das +3
Reasoning segmentation seeks pixel-accurate masks for targets referenced by complex, often implicit instructions, requiring context-dependent reasoning over the scene. Recent multi…
LBMamba: Locally Bi-directional Mamba
Jingwei Zhang, Xi Han, Hong Qin +2
Mamba, a State Space Model (SSM) that accelerates training by recasting recurrence as a parallel scan, has recently emerged as a linearly-scaling alternative to self-attention. Bec…
Region segmentation for sparse decompositions: better brain parcellations from rest fMRI
Alexandre Abraham, Elvis Dohmatob, Bertrand Thirion +2
Functional Magnetic Resonance Images acquired during resting-state provide information about the functional organization of the brain through measuring correlations between brain a…
Learning Probabilistic Topological Representations Using Discrete Morse Theory
Xiaoling Hu, Dimitris Samaras, Chao Chen
Accurate delineation of fine-scale structures is a very important yet challenging problem. Existing methods use topological information as an additional training loss, but are ulti…
OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following
Qiaomu Miao, Haoyu Wu, Jingyi Xu +2
Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure sp…
S-VolSDF: Sparse Multi-View Stereo Regularization of Neural Implicit Surfaces
Haoyu Wu, Alexandros Graikos, Dimitris Samaras
Neural rendering of implicit surfaces performs well in 3D vision applications. However, it requires dense input views as supervision. When only sparse input images are available, o…
Distribution Matching for Crowd Counting
Boyu Wang, Huidong Liu, Dimitris Samaras +1
In crowd counting, each training image contains multiple people, where each person is annotated by a dot. Existing crowd counting methods need to use a Gaussian to smooth each anno…
2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification
Jingwei Zhang, Anh Tien Nguyen, Xi Han +4
Efficiently modeling large 2D contexts is essential for various fields including Giga-Pixel Whole Slide Imaging (WSI) and remote sensing. Transformer-based models offer high parall…
Unsupervised Histopathology Image Synthesis
Le Hou, Ayush Agarwal, Dimitris Samaras +3
Hematoxylin and Eosin stained histopathology image analysis is essential for the diagnosis and study of complicated diseases such as cancer. Existing state-of-the-art approaches de…
Fast constrained sampling in pre-trained diffusion models
Alexandros Graikos, Nebojsa Jojic, Dimitris Samaras
Large denoising diffusion models, such as Stable Diffusion, have been trained on billions of image-caption pairs to perform text-conditioned image generation. As a byproduct of thi…
GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology
Saarthak Kapse, Pushpak Pati, Srikar Yellapragada +5
Pretraining a Multiple Instance Learning (MIL) aggregator enables the derivation of Whole Slide Image (WSI)-level embeddings from patch-level representations without supervision. W…
MI-NeRF: Learning a Single Face NeRF from Multiple Identities
Aggelina Chatziagapi, Grigorios G. Chrysos, Dimitris Samaras
In this work, we introduce a method that learns a single dynamic neural radiance field (NeRF) from monocular talking face videos of multiple identities. NeRFs have shown remarkable…
Learning 3D Reconstruction with Priors in Test Time
Lei Zhou, Haoyu Wu, Akshat Dave +1
We introduce a test-time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or…
Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning
Shilin Hu, Jingyi Xu, Sagnik Das +2
Shadows encode rich information about scene geometry and illumination, yet existing methods either predict a unified shadow mask or overlook attached shadows entirely. We address t…
Geodesic Distance Histogram Feature for Video Segmentation
Hieu Le, Vu Nguyen, Chen-Ping Yu +1
This paper proposes a geodesic-distance-based feature that encodes global information for improved video segmentation algorithms. The feature is a joint histogram of intensity and…
AS-Bridge: A Bidirectional Generative Framework Bridging Next-Generation Astronomical Surveys
Dichang Zhang, Yixuan Shao, Simon Birrer +1
The upcoming decade of observational cosmology will be shaped by large sky surveys, such as the ground-based LSST at the Vera C. Rubin Observatory and the space-based Euclid missio…
From Shadow Segmentation to Shadow Removal
Hieu Le, Dimitris Samaras
The requirement for paired shadow and shadow-free images limits the size and diversity of shadow removal datasets and hinders the possibility of training large-scale, robust shadow…
On the Statistical Efficiency of Multi-Task Learning of Gaussian Graphical Models
Jean Honorio, Tommi Jaakkola, Dimitris Samaras
In this paper, we present multi-task structure learning for Gaussian graphical models. We analyze the sufficient number of samples for the correct recovery of the supp…
Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes
Minh-Quan Le, Armand Comas, Alexandros Lattas +7
The paper proposes a self‑correcting framework called SC‑CMJP that couples image understanding and generation via cross‑modal attention in masked diffusion models, and introduces a…
Leveraging Registers in Vision Transformers for Robust Adaptation
Srikar Yellapragada, Kowshik Thopalli, Vivek Narayanaswamy +5
Vision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence o…
A Novel Framework for Characterization of Tumor-Immune Spatial Relationships in Tumor Microenvironment
Mahmudul Hasan, Jakub R. Kaczmarzyk, David Paredes +9
Understanding the impact of tumor biology on the composition of nearby cells often requires characterizing the impact of biologically distinct tumor regions. Biomarkers have been d…
Pathologist Attention-Aligned Report Generation for Prostate Histopathology
Ruoyu Xue, Suryakant Singh, Souradeep Chakraborty +12
The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracted from whole-slide images (W…
A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation
Hieu Le, Tomas F. Yago Vicente, Vu Nguyen +2
We propose a novel GAN-based framework for detecting shadows in images, in which a shadow detection network (D-Net) is trained together with a shadow attenuation network (A-Net) th…
LVLMs and Humans Ground Differently in Referential Communication
Peter Zeng, Weiling Li, Amie J. Paige +6
For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this ability to collaborate remains limited by a c…
-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions
Minh-Quan Le, Alexandros Graikos, Srikar Yellapragada +3
Synthesizing high-resolution images from intricate, domain-specific information remains a significant challenge in generative modeling, particularly for applications in large-image…
Dataset of Segmented Nuclei in Hematoxylin and Eosin Stained Histopathology Images of 10 Cancer Types
Le Hou, Rajarsi Gupta, John S. Van Arnam +5
The distribution and appearance of nuclei are essential markers for the diagnosis and study of cancer. Despite the importance of nuclear morphology, there is a lack of large scale,…
Deriving reproducible biomarkers from multi-site resting-state data: An Autism-based example
Alexandre Abraham, Michael Milham, Adriana Di Martino +4
Resting-state functional Magnetic Resonance Imaging (R-fMRI) holds the promise to reveal functional biomarkers of neuropsychiatric disorders. However, extracting such biomarkers is…
Attention De-sparsification Matters: Inducing Diversity in Digital Pathology Representation Learning
Saarthak Kapse, Srijan Das, Jingwei Zhang +4
We propose DiRL, a Diversity-inducing Representation Learning technique for histopathology imaging. Self-supervised learning techniques, such as contrastive and non-contrastive app…
SIDER: Single-Image Neural Optimization for Facial Geometric Detail Recovery
Aggelina Chatziagapi, ShahRukh Athar, Francesc Moreno-Noguer +1
We present SIDER(Single-Image neural optimization for facial geometric DEtail Recovery), a novel photometric optimization method that recovers detailed facial geometry from a singl…
Topo-R1: Detecting Topological Anomalies via Vision-Language Models
Meilong Xu, Qingqiao Hu, Xiaoling Hu +6
Topology is critical in tubular structures such as blood vessels, nerve fibers, and road networks, where connectivity and loop structure govern downstream functional analysis. Visi…
GeoMaNO: Geometric Mamba Neural Operator for Partial Differential Equations
Xi Han, Jingwei Zhang, Dimitris Samaras +2
The neural operator (NO) framework has emerged as a powerful tool for solving partial differential equations (PDEs). Recent NOs are dominated by the Transformer architecture, which…
Patch-level Gaze Distribution Prediction for Gaze Following
Qiaomu Miao, Minh Hoai, Dimitris Samaras
Gaze following aims to predict where a person is looking in a scene, by predicting the target location, or indicating that the target is located outside the image. Recent works det…
Zero-shot Object Counting
Jingyi Xu, Hieu Le, Vu Nguyen +2
Class-agnostic object counting aims to count object instances of an arbitrary class at test time. It is challenging but also enables many potential applications. Current methods re…
Co-localization with Category-Consistent Features and Geodesic Distance Propagation
Hieu Le, Chen-Ping Yu, Gregory Zelinsky +1
Co-localization is the problem of localizing objects of the same class using only the set of images that contain them. This is a challenging task because the object detector must b…
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
Kostas Triaridis, Alexandros Graikos, Aggelina Chatziagapi +2
Hallucinations in diffusion models are samples with structural inconsistencies that can emerge due to the excessive smoothing of the learned score function, which in turn leads to…
Gigapixel Whole-Slide Images Classification using Locally Supervised Learning
Jingwei Zhang, Xin Zhang, Ke Ma +4
Histopathology whole slide images (WSIs) play a very important role in clinical studies and serve as the gold standard for many cancer diagnoses. However, generating automatic tool…
Neural Networks with Smooth Adaptive Activation Functions for Regression
Le Hou, Dimitris Samaras, Tahsin M. Kurc +2
In Neural Networks (NN), Adaptive Activation Functions (AAF) have parameters that control the shapes of activation functions. These parameters are trained along with other paramete…
Zero-Shot Object Counting with Language-Vision Models
Jingyi Xu, Hieu Le, Dimitris Samaras
Class-agnostic object counting aims to count object instances of an arbitrary class at test time. It is challenging but also enables many potential applications. Current methods re…
Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers
Zhibo Yang, Sounak Mondal, Seoyoung Ahn +4
Most models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose th…
Decoding the visual attention of pathologists to reveal their level of expertise
Souradeep Chakraborty, Dana Perez, Paul Friedman +7
We present a method for classifying the expertise of a pathologist based on how they allocated their attention during a cancer reading. We engage this decoding task by developing a…
MORCIC: Model Order Reduction Techniques for Electromagnetic Models of Integrated Circuits
Dimitrios Garyfallou, Athanasios Stefanou, Christos Giamouzis +14
Model order reduction (MOR) is crucial for the design process of integrated circuits. Specifically, the vast amount of passive RLCk elements in electromagnetic models extracted fro…
Gen-SIS: Generative Self-augmentation Improves Self-supervised Learning
Varun Belagali, Srikar Yellapragada, Alexandros Graikos +7
Self-supervised learning (SSL) methods have emerged as strong visual representation learners by training an image encoder to maximize similarity between features of different views…
TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans
Aggelina Chatziagapi, Bindita Chaudhuri, Amit Kumar +3
We introduce a novel framework that learns a dynamic neural radiance field (NeRF) for full-body talking humans from monocular videos. Prior work represents only the body pose or th…
Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks
Le Hou, Chen-Ping Yu, Dimitris Samaras
In the context of single-label classification, despite the huge success of deep learning, the commonly used cross-entropy loss function ignores the intricate inter-class relationsh…
GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation
Snehal Singh Tomar, Alexandros Graikos, Arjun Krishna +2
Modern deep learning methods typically treat image sequences as large tensors of sequentially stacked frames. However, is this straightforward representation ideal given the curren…
Shadow Removal via Shadow Image Decomposition
Hieu Le, Dimitris Samaras
We propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects…
Target-absent Human Attention
Zhibo Yang, Sounak Mondal, Seoyoung Ahn +3
The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user's attention. Computer vision models have been develope…
Physics-based Shadow Image Decomposition for Shadow Removal
Hieu Le, Dimitris Samaras
We propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects…
An Adversarial Neuro-Tensorial Approach For Learning Disentangled Representations
Mengjiao Wang, Zhixin Shu, Shiyang Cheng +3
Several factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of varia…
Self-supervised co-salient object detection via feature correspondence at multiple scales
Souradeep Chakraborty, Dimitris Samaras
Our paper introduces a novel two-stage self-supervised approach for detecting co-occurring salient objects (CoSOD) in image groups without requiring segmentation annotations. Unlik…
Weighting Pseudo-Labels via High-Activation Feature Index Similarity and Object Detection for Semi-Supervised Segmentation
Prantik Howlader, Hieu Le, Dimitris Samaras
Semi-supervised semantic segmentation methods leverage unlabeled data by pseudo-labeling them. Thus the success of these methods hinges on the reliablility of the pseudo-labels. Ex…
Multi-Class Cell Detection Using Spatial Context Representation
Shahira Abousamra, David Belinsky, John Van Arnam +7
In digital pathology, both detection and classification of cells are important for automatic diagnostic and prognostic tasks. Classifying cells into subtypes, such as tumor cells,…
ZoomLDM: Latent Diffusion Model for multi-scale image generation
Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis +4
Diffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Gi…
Token Sparsification for Faster Medical Image Segmentation
Lei Zhou, Huidong Liu, Joseph Bae +3
Can we use sparse tokens for dense prediction, e.g., segmentation? Although token sparsification has been applied to Vision Transformers (ViT) to accelerate classification, it is s…
TopoCellGen: Generating Histopathology Cell Topology with a Diffusion Model
Meilong Xu, Saumya Gupta, Xiaoling Hu +5
Accurately modeling multi-class cell topology is crucial in digital pathology, as it provides critical insights into tissue structure and pathology. The synthetic generation of cel…
Conditional Generation from Unconditional Diffusion Models using Denoiser Representations
Alexandros Graikos, Srikar Yellapragada, Dimitris Samaras
Denoising diffusion models have gained popularity as a generative modeling technique for producing high-quality and diverse images. Applying these models to downstream tasks requir…
Local Learning on Transformers via Feature Reconstruction
Priyank Pathak, Jingwei Zhang, Dimitris Samaras
Transformers are becoming increasingly popular due to their superior performance over conventional convolutional neural networks(CNNs). However, transformers usually require a much…
Label Super Resolution with Inter-Instance Loss
Maozheng Zhao, Le Hou, Han Le +8
For the task of semantic segmentation, high-resolution (pixel-level) ground truth is very expensive to collect, especially for high resolution images such as gigapixel pathology im…
A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies
Panagiotis Kaliosis, Adithya V Ganesan, Oscar N. E. Kjell +8
Large language models (LLMs) are increasingly being used in a zero-shot (generative) fashion to assess mental health conditions, yet we have limited knowledge on what factors affec…
Utilizing Automated Breast Cancer Detection to Identify Spatial Distributions of Tumor Infiltrating Lymphocytes in Invasive Breast Cancer
Han Le, Rajarsi Gupta, Le Hou +12
Quantitative assessment of Tumor-TIL spatial relationships is increasingly important in both basic science and clinical aspects of breast cancer research. We have developed and eva…
Phrase-Instance Alignment for Generalized Referring Segmentation
E-Ro Nguyen, Hieu Le, Dimitris Samaras +1
Generalized Referring expressions can describe one object, several related objects, or none at all. Existing generalized referring segmentation (GRES) models treat all cases alike,…
Prompt-MIL: Boosting Multi-Instance Learning Schemes via Task-specific Prompt Tuning
Jingwei Zhang, Saarthak Kapse, Ke Ma +4
Whole slide image (WSI) classification is a critical task in computational pathology, requiring the processing of gigapixel-sized images, which is challenging for current deep-lear…
Learned representation-guided diffusion models for large-image generation
Alexandros Graikos, Srikar Yellapragada, Minh-Quan Le +4
To synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-…
Self-supervised Deformation Modeling for Facial Expression Editing
ShahRukh Athar, Zhixin Shu, Dimitris Samaras
Recent advances in deep generative models have demonstrated impressive results in photo-realistic facial image synthesis and editing. Facial expressions are inherently the result o…
Patch-based Convolutional Neural Network for Whole Slide Tissue Image Classification
Le Hou, Dimitris Samaras, Tahsin M. Kurc +3
Convolutional Neural Networks (CNN) are state-of-the-art models for many image classification tasks. However, to recognize cancer subtypes automatically, training a CNN on gigapixe…
Topology-Guided Multi-Class Cell Context Generation for Digital Pathology
Shahira Abousamra, Rajarsi Gupta, Tahsin Kurc +3
In digital pathology, the spatial context of cells is important for cell classification, cancer diagnosis and prognosis. To model such complex cell context, however, is challenging…
Topology-Aware Segmentation Using Discrete Morse Theory
Xiaoling Hu, Yusu Wang, Li Fuxin +2
In the segmentation of fine-scale structures from natural and biomedical images, per-pixel accuracy is not the only metric of concern. Topological correctness, such as vessel conne…
Latent Space Optimal Transport for Generative Models
Huidong Liu, Yang Guo, Na Lei +4
Variational Auto-Encoders enforce their learned intermediate latent-space data distribution to be a simple distribution, such as an isotropic Gaussian. However, this causes the pos…
Center-Focusing Multi-task CNN with Injected Features for Classification of Glioma Nuclear Images
Veda Murthy, Le Hou, Dimitris Samaras +2
Classifying the various shapes and attributes of a glioma cell nucleus is crucial for diagnosis and understanding the disease. We investigate automated classification of glioma nuc…
Learning Relighting and Intrinsic Decomposition in Neural Radiance Fields
Yixiong Yang, Shilin Hu, Haoyu Wu +3
The task of extracting intrinsic components, such as reflectance and shading, from neural radiance fields is of growing interest. However, current methods largely focus on syntheti…
A Study of Human Gaze Behavior During Visual Crowd Counting
Raji Annadi, Yupei Chen, Viresh Ranjan +3
In this paper, we describe our study on how humans allocate their attention during visual crowd counting. Using an eye tracker, we collect gaze behavior of human participants who a…
Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning
Zhibo Yang, Lihan Huang, Yupei Chen +5
Being able to predict human gaze behavior has obvious importance for behavioral vision and for computer vision applications. Most models have mainly focused on predicting free-view…
Toward ultra-efficient high fidelity predictions of wind turbine wakes: Augmenting the accuracy of engineering models via LES-trained machine learning
Christian Santoni, Dichang Zhang, Zexia Zhang +3
This study proposes a novel machine learning (ML) methodology for the efficient and cost-effective prediction of high-fidelity three-dimensional velocity fields in the wake of util…
Localization in the Crowd with Topological Constraints
Shahira Abousamra, Minh Hoai, Dimitris Samaras +1
We address the problem of crowd localization, i.e., the prediction of dots corresponding to people in a crowded scene. Due to various challenges, a localization method is prone to…
Topology-Preserving Deep Image Segmentation
Xiaoling Hu, Li Fuxin, Dimitris Samaras +1
Segmentation algorithms are prone to make topological errors on fine-scale structures, e.g., broken connections. We propose a novel method that learns to segment with correct topol…
Intrinsic Decomposition of Document Images In-the-Wild
Sagnik Das, Hassan Ahmed Sial, Ke Ma +3
Automatic document content processing is affected by artifacts caused by the shape of the paper, non-uniform and diverse color of lighting conditions. Fully-supervised methods on r…
CDG-MAE: Cross-view Masked Modeling using Diffusion Generated Views
Varun Belagali, Pierre Marza, Srikar Yellapragada +7
Cross-view masked autoencoding has emerged as a powerful pretext task for learning dense correspondences, which are essential for applications such as video label propagation. The…
Learning from Thresholds: Fully Automated Classification of Tumor Infiltrating Lymphocytes for Multiple Cancer Types
Shahira Abousamra, Le Hou, Rajarsi Gupta +7
Deep learning classifiers for characterization of whole slide tissue morphology require large volumes of annotated data to learn variations across different tissue and cancer types…
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
Qiaomu Miao, Alexandros Graikos, Jingwei Zhang +3
Training gaze following models requires a large number of images with gaze target coordinates annotated by human annotators, which is a laborious and inherently ambiguous process.…
Unsupervised and semi-supervised co-salient object detection via segmentation frequency statistics
Souradeep Chakraborty, Shujon Naha, Muhammet Bastan +2
In this paper, we address the detection of co-occurring salient objects (CoSOD) in an image group using frequency statistics in an unsupervised manner, which further enable us to d…
Weakly Labeling the Antarctic: The Penguin Colony Case
Hieu Le, Bento Gonçalves, Dimitris Samaras +1
Antarctic penguins are important ecological indicators -- especially in the face of climate change. In this work, we present a deep learning based model for semantic segmentation o…
Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
Sounak Mondal, Dimitris Samaras, Gregory Zelinsky +1
We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as…
What is the Right Embedding Space for Contrastive Learning in REC?
Kostas Triaridis, Panagiotis Kaliosis, E-Ro Nguyen +3
Referring Expression Counting (REC) requires distinguishing visually similar objects described by fine-grained text cues. Existing methods tackle this via image-text contrastive le…
PathSegDiff: Pathology Segmentation using Diffusion model representations
Sachin Kumar Danisetty, Alexandros Graikos, Srikar Yellapragada +1
Image segmentation is crucial in many computational pathology pipelines, including accurate disease diagnosis, subtyping, outcome, and survivability prediction. The common approach…
Computational Pathology: A Survey Review and The Way Forward
Mahdi S. Hosseini, Babak Ehteshami Bejnordi, Vincent Quoc-Huy Trinh +18
Computational Pathology CPath is an interdisciplinary science that augments developments of computational approaches to analyze and model medical histopathology images. The main ob…
Light Direction and Color Estimation from Single Image with Deep Regression
Hassan A. Sial, Ramon Baldrich, Maria Vanrell +1
We present a method to estimate the direction and color of the scene light source from a single image. Our method is based on two main ideas: (a) we use a new synthetic dataset wit…
Modeling Deep Learning Based Privacy Attacks on Physical Mail
Bingyao Huang, Ruyi Lian, Dimitris Samaras +1
Mail privacy protection aims to prevent unauthorized access to hidden content within an envelope since normal paper envelopes are not as safe as we think. In this paper, for the fi…
MLI-NeRF: Multi-Light Intrinsic-Aware Neural Radiance Fields
Yixiong Yang, Shilin Hu, Haoyu Wu +3
Current methods for extracting intrinsic image components, such as reflectance and shading, primarily rely on statistical priors. These methods focus mainly on simple synthetic sce…
Counting Trees from Satellite Imagery with Noisy Supervision
Dimitri Gominski, Maurice Mugabowindekwe, Qiue Xu +6
Counting individual trees is a fundamental task for environmental monitoring, yet remains largely unexplored with satellite imagery. At these resolutions, isolated trees may still…
FaceDet3D: Facial Expressions with 3D Geometric Detail Prediction
ShahRukh Athar, Albert Pumarola, Francesc Moreno-Noguer +1
Facial Expressions induce a variety of high-level details on the 3D face geometry. For example, a smile causes the wrinkling of cheeks or the formation of dimples, while being angr…
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
Minh-Quan Le, Gaurav Mittal, Cheng Zhao +3
Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training…
Exascale Deep Learning to Accelerate Cancer Research
Robert M. Patton, J. Travis Johnston, Steven R. Young +9
Deep learning, through the use of neural networks, has demonstrated remarkable ability to automate many routine tasks when presented with sufficient data for training. The neural n…
Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer
Haoyu Wu, Jingyi Xu, Qiaomu Miao +2
Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-resolution tokens remains undere…
Few-shot Personalized Scanpath Prediction
Ruoyu Xue, Jingyi Xu, Sounak Mondal +4
A personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training sca…