papers

Publications (152)

cs.CV2026

Embedding Physical Reasoning into Diffusion-Based Shadow Generation

Shilin Hu, Jingyi Xu, Akshat Dave +2

Generating realistic shadows for inserted objects requires reasoning about scene geometry and illumination. However, most existing methods operate purely in image space, leaving th…

cs.CV2025

Talking Head Generation via AU-Guided Landmark Prediction

Shao-Yu Chang, Jingyi Xu, Hieu Le +1

We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotio…

cs.CV2025

Personalized Image Descriptions from Attention Sequences

Ruoyu Xue, Hieu Le, Jingyi Xu +5

People can view the same image differently: they focus on different regions, objects, and details in varying orders and describe them in distinct linguistic styles. This leads to s…

cs.CV2024

Predicting Visual Attention in Graphic Design Documents

Souradeep Chakraborty, Zijun Wei, Conor Kelton +4

We present a model for predicting visual attention during the free viewing of graphic design documents. While existing works on this topic have aimed at predicting static saliency…

cs.CV2024

SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel Histopathology

Saarthak Kapse, Pushpak Pati, Srijan Das +7

Introducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging, given the complexity of gigapixel slid…

cs.CV2026

Generating metamers of human scene understanding

Ritik Raina, Abe Leite, Alexandros Graikos +3

Human vision combines low-resolution "gist" information from the visual periphery with sparse but high-resolution information from fixated locations to construct a coherent underst…

cs.CV2025

CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation

Prantik Howlader, Hoang Nguyen-Canh, Srijan Das +3

Reasoning segmentation seeks pixel-accurate masks for targets referenced by complex, often implicit instructions, requiring context-dependent reasoning over the scene. Recent multi…

cs.CV2025

LBMamba: Locally Bi-directional Mamba

Jingwei Zhang, Xi Han, Hong Qin +2

Mamba, a State Space Model (SSM) that accelerates training by recasting recurrence as a parallel scan, has recently emerged as a linearly-scaling alternative to self-attention. Bec…

q-bio.NC2014

Region segmentation for sparse decompositions: better brain parcellations from rest fMRI

Alexandre Abraham, Elvis Dohmatob, Bertrand Thirion +2

Functional Magnetic Resonance Images acquired during resting-state provide information about the functional organization of the brain through measuring correlations between brain a…

eess.IV2022

Learning Probabilistic Topological Representations Using Discrete Morse Theory

Xiaoling Hu, Dimitris Samaras, Chao Chen

Accurate delineation of fine-scale structures is a very important yet challenging problem. Existing methods use topological information as an additional training loss, but are ulti…

cs.CV2026

OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following

Qiaomu Miao, Haoyu Wu, Jingyi Xu +2

Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typically restricted to pure sp…

cs.CV2023

S-VolSDF: Sparse Multi-View Stereo Regularization of Neural Implicit Surfaces

Haoyu Wu, Alexandros Graikos, Dimitris Samaras

Neural rendering of implicit surfaces performs well in 3D vision applications. However, it requires dense input views as supervision. When only sparse input images are available, o…

cs.CV2020

Distribution Matching for Crowd Counting

Boyu Wang, Huidong Liu, Dimitris Samaras +1

In crowd counting, each training image contains multiple people, where each person is annotated by a dot. Existing crowd counting methods need to use a Gaussian to smooth each anno…

cs.CV2026

2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification

Jingwei Zhang, Anh Tien Nguyen, Xi Han +4

Efficiently modeling large 2D contexts is essential for various fields including Giga-Pixel Whole Slide Imaging (WSI) and remote sensing. Transformer-based models offer high parall…

cs.CV2017

Unsupervised Histopathology Image Synthesis

Le Hou, Ayush Agarwal, Dimitris Samaras +3

Hematoxylin and Eosin stained histopathology image analysis is essential for the diagnosis and study of complicated diseases such as cancer. Existing state-of-the-art approaches de…

cs.CV2025

Fast constrained sampling in pre-trained diffusion models

Alexandros Graikos, Nebojsa Jojic, Dimitris Samaras

Large denoising diffusion models, such as Stable Diffusion, have been trained on billions of image-caption pairs to perform text-conditioned image generation. As a byproduct of thi…

cs.CV2025

GECKO: Gigapixel Vision-Concept Contrastive Pretraining in Histopathology

Saarthak Kapse, Pushpak Pati, Srikar Yellapragada +5

Pretraining a Multiple Instance Learning (MIL) aggregator enables the derivation of Whole Slide Image (WSI)-level embeddings from patch-level representations without supervision. W…

cs.CV2024

MI-NeRF: Learning a Single Face NeRF from Multiple Identities

Aggelina Chatziagapi, Grigorios G. Chrysos, Dimitris Samaras

In this work, we introduce a method that learns a single dynamic neural radiance field (NeRF) from monocular talking face videos of multiple identities. NeRFs have shown remarkable…

cs.CV2026

Learning 3D Reconstruction with Priors in Test Time

Lei Zhou, Haoyu Wu, Akshat Dave +1

We introduce a test-time framework for multiview Transformers (MVTs) that incorporates priors (e.g., camera poses, intrinsics, and depth) to improve 3D tasks without retraining or…

cs.CV2026

Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning

Shilin Hu, Jingyi Xu, Sagnik Das +2

Shadows encode rich information about scene geometry and illumination, yet existing methods either predict a unified shadow mask or overlook attached shadows entirely. We address t…

cs.CV2017

Geodesic Distance Histogram Feature for Video Segmentation

Hieu Le, Vu Nguyen, Chen-Ping Yu +1

This paper proposes a geodesic-distance-based feature that encodes global information for improved video segmentation algorithms. The feature is a joint histogram of intensity and…

astro-ph.IM2026

AS-Bridge: A Bidirectional Generative Framework Bridging Next-Generation Astronomical Surveys

Dichang Zhang, Yixuan Shao, Simon Birrer +1

The upcoming decade of observational cosmology will be shaped by large sky surveys, such as the ground-based LSST at the Vera C. Rubin Observatory and the space-based Euclid missio…

cs.CV2020

From Shadow Segmentation to Shadow Removal

Hieu Le, Dimitris Samaras

The requirement for paired shadow and shadow-free images limits the size and diversity of shadow removal datasets and hinders the possibility of training large-scale, robust shadow…

cs.LG2015

On the Statistical Efficiency of Multi-Task Learning of Gaussian Graphical Models

Jean Honorio, Tommi Jaakkola, Dimitris Samaras

In this paper, we present multi-task structure learning for Gaussian graphical models. We analyze the sufficient number of samples for the correct recovery of the supp…

cs.LG2026

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes

Minh-Quan Le, Armand Comas, Alexandros Lattas +7

The paper proposes a self‑correcting framework called SC‑CMJP that couples image understanding and generation via cross‑modal attention in masked diffusion models, and introduces a…

#multimodal generation#image understanding#masked diffusion models#cross‑modal attention
cs.CV2025

Leveraging Registers in Vision Transformers for Robust Adaptation

Srikar Yellapragada, Kowshik Thopalli, Vivek Narayanaswamy +5

Vision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence o…

q-bio.QM2022

A Novel Framework for Characterization of Tumor-Immune Spatial Relationships in Tumor Microenvironment

Mahmudul Hasan, Jakub R. Kaczmarzyk, David Paredes +9

Understanding the impact of tumor biology on the composition of nearby cells often requires characterizing the impact of biologically distinct tumor regions. Biomarkers have been d…

cs.CV2026

Pathologist Attention-Aligned Report Generation for Prostate Histopathology

Ruoyu Xue, Suryakant Singh, Souradeep Chakraborty +12

The allocation of visual attention by pathologists during cancer diagnosis is a highly selective process that critically shapes the information extracted from whole-slide images (W…

cs.CV2018

A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation

Hieu Le, Tomas F. Yago Vicente, Vu Nguyen +2

We propose a novel GAN-based framework for detecting shadows in images, in which a shadow detection network (D-Net) is trained together with a shadow attenuation network (A-Net) th…

cs.CL2026

LVLMs and Humans Ground Differently in Referential Communication

Peter Zeng, Weiling Li, Amie J. Paige +6

For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this ability to collaborate remains limited by a c…

cs.CV2024

-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions

Minh-Quan Le, Alexandros Graikos, Srikar Yellapragada +3

Synthesizing high-resolution images from intricate, domain-specific information remains a significant challenge in generative modeling, particularly for applications in large-image…

eess.IV2020

Dataset of Segmented Nuclei in Hematoxylin and Eosin Stained Histopathology Images of 10 Cancer Types

Le Hou, Rajarsi Gupta, John S. Van Arnam +5

The distribution and appearance of nuclei are essential markers for the diagnosis and study of cancer. Despite the importance of nuclear morphology, there is a lack of large scale,…

stat.ML2016

Deriving reproducible biomarkers from multi-site resting-state data: An Autism-based example

Alexandre Abraham, Michael Milham, Adriana Di Martino +4

Resting-state functional Magnetic Resonance Imaging (R-fMRI) holds the promise to reveal functional biomarkers of neuropsychiatric disorders. However, extracting such biomarkers is…

cs.CV2023

Attention De-sparsification Matters: Inducing Diversity in Digital Pathology Representation Learning

Saarthak Kapse, Srijan Das, Jingwei Zhang +4

We propose DiRL, a Diversity-inducing Representation Learning technique for histopathology imaging. Self-supervised learning techniques, such as contrastive and non-contrastive app…

cs.CV2021

SIDER: Single-Image Neural Optimization for Facial Geometric Detail Recovery

Aggelina Chatziagapi, ShahRukh Athar, Francesc Moreno-Noguer +1

We present SIDER(Single-Image neural optimization for facial geometric DEtail Recovery), a novel photometric optimization method that recovers detailed facial geometry from a singl…

cs.CV2026

Topo-R1: Detecting Topological Anomalies via Vision-Language Models

Meilong Xu, Qingqiao Hu, Xiaoling Hu +6

Topology is critical in tubular structures such as blood vessels, nerve fibers, and road networks, where connectivity and loop structure govern downstream functional analysis. Visi…

cs.LG2025

GeoMaNO: Geometric Mamba Neural Operator for Partial Differential Equations

Xi Han, Jingwei Zhang, Dimitris Samaras +2

The neural operator (NO) framework has emerged as a powerful tool for solving partial differential equations (PDEs). Recent NOs are dominated by the Transformer architecture, which…

cs.CV2022

Patch-level Gaze Distribution Prediction for Gaze Following

Qiaomu Miao, Minh Hoai, Dimitris Samaras

Gaze following aims to predict where a person is looking in a scene, by predicting the target location, or indicating that the target is located outside the image. Recent works det…

cs.CV2023

Zero-shot Object Counting

Jingyi Xu, Hieu Le, Vu Nguyen +2

Class-agnostic object counting aims to count object instances of an arbitrary class at test time. It is challenging but also enables many potential applications. Current methods re…

cs.CV2019

Co-localization with Category-Consistent Features and Geodesic Distance Propagation

Hieu Le, Chen-Ping Yu, Gregory Zelinsky +1

Co-localization is the problem of localizing objects of the same class using only the set of images that contain them. This is a challenging task because the object detector must b…

cs.CV2026

Mitigating Diffusion Model Hallucinations with Dynamic Guidance

Kostas Triaridis, Alexandros Graikos, Aggelina Chatziagapi +2

Hallucinations in diffusion models are samples with structural inconsistencies that can emerge due to the excessive smoothing of the learned score function, which in turn leads to…

cs.CV2022

Gigapixel Whole-Slide Images Classification using Locally Supervised Learning

Jingwei Zhang, Xin Zhang, Ke Ma +4

Histopathology whole slide images (WSIs) play a very important role in clinical studies and serve as the gold standard for many cancer diagnoses. However, generating automatic tool…

cs.CV2016

Neural Networks with Smooth Adaptive Activation Functions for Regression

Le Hou, Dimitris Samaras, Tahsin M. Kurc +2

In Neural Networks (NN), Adaptive Activation Functions (AAF) have parameters that control the shapes of activation functions. These parameters are trained along with other paramete…

cs.CV2023

Zero-Shot Object Counting with Language-Vision Models

Jingyi Xu, Hieu Le, Dimitris Samaras

Class-agnostic object counting aims to count object instances of an arbitrary class at test time. It is challenging but also enables many potential applications. Current methods re…

cs.CV2024

Unifying Top-down and Bottom-up Scanpath Prediction Using Transformers

Zhibo Yang, Sounak Mondal, Seoyoung Ahn +4

Most models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose th…

eess.IV2024

Decoding the visual attention of pathologists to reveal their level of expertise

Souradeep Chakraborty, Dana Perez, Paul Friedman +7

We present a method for classifying the expertise of a pathologist based on how they allocated their attention during a cancer reading. We engage this decoding task by developing a…

math.NA2023

MORCIC: Model Order Reduction Techniques for Electromagnetic Models of Integrated Circuits

Dimitrios Garyfallou, Athanasios Stefanou, Christos Giamouzis +14

Model order reduction (MOR) is crucial for the design process of integrated circuits. Specifically, the vast amount of passive RLCk elements in electromagnetic models extracted fro…

cs.CV2024

Gen-SIS: Generative Self-augmentation Improves Self-supervised Learning

Varun Belagali, Srikar Yellapragada, Alexandros Graikos +7

Self-supervised learning (SSL) methods have emerged as strong visual representation learners by training an image encoder to maximize similarity between features of different views…

cs.CV2024

TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans

Aggelina Chatziagapi, Bindita Chaudhuri, Amit Kumar +3

We introduce a novel framework that learns a dynamic neural radiance field (NeRF) for full-body talking humans from monocular videos. Prior work represents only the body pose or th…

cs.CV2017

Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks

Le Hou, Chen-Ping Yu, Dimitris Samaras

In the context of single-label classification, despite the huge success of deep learning, the commonly used cross-entropy loss function ignores the intricate inter-class relationsh…

cs.CV2026

GriDiT: Factorized Grid-Based Diffusion for Efficient Long Image Sequence Generation

Snehal Singh Tomar, Alexandros Graikos, Arjun Krishna +2

Modern deep learning methods typically treat image sequences as large tensors of sequentially stacked frames. However, is this straightforward representation ideal given the curren…

cs.CV2019

Shadow Removal via Shadow Image Decomposition

Hieu Le, Dimitris Samaras

We propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects…

cs.CV2022

Target-absent Human Attention

Zhibo Yang, Sounak Mondal, Seoyoung Ahn +3

The prediction of human gaze behavior is important for building human-computer interactive systems that can anticipate a user's attention. Computer vision models have been develope…

cs.CV2021

Physics-based Shadow Image Decomposition for Shadow Removal

Hieu Le, Dimitris Samaras

We propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects…

cs.CV2018

An Adversarial Neuro-Tensorial Approach For Learning Disentangled Representations

Mengjiao Wang, Zhixin Shu, Shiyang Cheng +3

Several factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of varia…

cs.CV2024

Self-supervised co-salient object detection via feature correspondence at multiple scales

Souradeep Chakraborty, Dimitris Samaras

Our paper introduces a novel two-stage self-supervised approach for detecting co-occurring salient objects (CoSOD) in image groups without requiring segmentation annotations. Unlik…

cs.CV2024

Weighting Pseudo-Labels via High-Activation Feature Index Similarity and Object Detection for Semi-Supervised Segmentation

Prantik Howlader, Hieu Le, Dimitris Samaras

Semi-supervised semantic segmentation methods leverage unlabeled data by pseudo-labeling them. Thus the success of these methods hinges on the reliablility of the pseudo-labels. Ex…

cs.CV2022

Multi-Class Cell Detection Using Spatial Context Representation

Shahira Abousamra, David Belinsky, John Van Arnam +7

In digital pathology, both detection and classification of cells are important for automatic diagnostic and prognostic tasks. Classifying cells into subtypes, such as tumor cells,…

cs.CV2025

ZoomLDM: Latent Diffusion Model for multi-scale image generation

Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis +4

Diffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Gi…

cs.CV2023

Token Sparsification for Faster Medical Image Segmentation

Lei Zhou, Huidong Liu, Joseph Bae +3

Can we use sparse tokens for dense prediction, e.g., segmentation? Although token sparsification has been applied to Vision Transformers (ViT) to accelerate classification, it is s…

eess.IV2025

TopoCellGen: Generating Histopathology Cell Topology with a Diffusion Model

Meilong Xu, Saumya Gupta, Xiaoling Hu +5

Accurately modeling multi-class cell topology is crucial in digital pathology, as it provides critical insights into tissue structure and pathology. The synthetic generation of cel…

cs.CV2023

Conditional Generation from Unconditional Diffusion Models using Denoiser Representations

Alexandros Graikos, Srikar Yellapragada, Dimitris Samaras

Denoising diffusion models have gained popularity as a generative modeling technique for producing high-quality and diverse images. Applying these models to downstream tasks requir…

cs.CV2022

Local Learning on Transformers via Feature Reconstruction

Priyank Pathak, Jingwei Zhang, Dimitris Samaras

Transformers are becoming increasingly popular due to their superior performance over conventional convolutional neural networks(CNNs). However, transformers usually require a much…

cs.CV2020

Label Super Resolution with Inter-Instance Loss

Maozheng Zhao, Le Hou, Han Le +8

For the task of semantic segmentation, high-resolution (pixel-level) ground truth is very expensive to collect, especially for high resolution images such as gigapixel pathology im…

cs.CL2026

A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies

Panagiotis Kaliosis, Adithya V Ganesan, Oscar N. E. Kjell +8

Large language models (LLMs) are increasingly being used in a zero-shot (generative) fashion to assess mental health conditions, yet we have limited knowledge on what factors affec…

eess.IV2020

Utilizing Automated Breast Cancer Detection to Identify Spatial Distributions of Tumor Infiltrating Lymphocytes in Invasive Breast Cancer

Han Le, Rajarsi Gupta, Le Hou +12

Quantitative assessment of Tumor-TIL spatial relationships is increasingly important in both basic science and clinical aspects of breast cancer research. We have developed and eva…

cs.CV2026

Phrase-Instance Alignment for Generalized Referring Segmentation

E-Ro Nguyen, Hieu Le, Dimitris Samaras +1

Generalized Referring expressions can describe one object, several related objects, or none at all. Existing generalized referring segmentation (GRES) models treat all cases alike,…

cs.CV2023

Prompt-MIL: Boosting Multi-Instance Learning Schemes via Task-specific Prompt Tuning

Jingwei Zhang, Saarthak Kapse, Ke Ma +4

Whole slide image (WSI) classification is a critical task in computational pathology, requiring the processing of gigapixel-sized images, which is challenging for current deep-lear…

cs.CV2024

Learned representation-guided diffusion models for large-image generation

Alexandros Graikos, Srikar Yellapragada, Minh-Quan Le +4

To synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-…

cs.CV2019

Self-supervised Deformation Modeling for Facial Expression Editing

ShahRukh Athar, Zhixin Shu, Dimitris Samaras

Recent advances in deep generative models have demonstrated impressive results in photo-realistic facial image synthesis and editing. Facial expressions are inherently the result o…

cs.CV2016

Patch-based Convolutional Neural Network for Whole Slide Tissue Image Classification

Le Hou, Dimitris Samaras, Tahsin M. Kurc +3

Convolutional Neural Networks (CNN) are state-of-the-art models for many image classification tasks. However, to recognize cancer subtypes automatically, training a CNN on gigapixe…

eess.IV2023

Topology-Guided Multi-Class Cell Context Generation for Digital Pathology

Shahira Abousamra, Rajarsi Gupta, Tahsin Kurc +3

In digital pathology, the spatial context of cells is important for cell classification, cancer diagnosis and prognosis. To model such complex cell context, however, is challenging…

cs.CV2021

Topology-Aware Segmentation Using Discrete Morse Theory

Xiaoling Hu, Yusu Wang, Li Fuxin +2

In the segmentation of fine-scale structures from natural and biomedical images, per-pixel accuracy is not the only metric of concern. Topological correctness, such as vessel conne…

cs.LG2018

Latent Space Optimal Transport for Generative Models

Huidong Liu, Yang Guo, Na Lei +4

Variational Auto-Encoders enforce their learned intermediate latent-space data distribution to be a simple distribution, such as an isotropic Gaussian. However, this causes the pos…

cs.CV2017

Center-Focusing Multi-task CNN with Injected Features for Classification of Glioma Nuclear Images

Veda Murthy, Le Hou, Dimitris Samaras +2

Classifying the various shapes and attributes of a glioma cell nucleus is crucial for diagnosis and understanding the disease. We investigate automated classification of glioma nuc…

cs.CV2024

Learning Relighting and Intrinsic Decomposition in Neural Radiance Fields

Yixiong Yang, Shilin Hu, Haoyu Wu +3

The task of extracting intrinsic components, such as reflectance and shading, from neural radiance fields is of growing interest. However, current methods largely focus on syntheti…

cs.CV2020

A Study of Human Gaze Behavior During Visual Crowd Counting

Raji Annadi, Yupei Chen, Viresh Ranjan +3

In this paper, we describe our study on how humans allocate their attention during visual crowd counting. Using an eye tracker, we collect gaze behavior of human participants who a…

cs.CV2020

Predicting Goal-directed Human Attention Using Inverse Reinforcement Learning

Zhibo Yang, Lihan Huang, Yupei Chen +5

Being able to predict human gaze behavior has obvious importance for behavioral vision and for computer vision applications. Most models have mainly focused on predicting free-view…

physics.flu-dyn2024

Toward ultra-efficient high fidelity predictions of wind turbine wakes: Augmenting the accuracy of engineering models via LES-trained machine learning

Christian Santoni, Dichang Zhang, Zexia Zhang +3

This study proposes a novel machine learning (ML) methodology for the efficient and cost-effective prediction of high-fidelity three-dimensional velocity fields in the wake of util…

cs.CV2020

Localization in the Crowd with Topological Constraints

Shahira Abousamra, Minh Hoai, Dimitris Samaras +1

We address the problem of crowd localization, i.e., the prediction of dots corresponding to people in a crowded scene. Due to various challenges, a localization method is prone to…

cs.CV2019

Topology-Preserving Deep Image Segmentation

Xiaoling Hu, Li Fuxin, Dimitris Samaras +1

Segmentation algorithms are prone to make topological errors on fine-scale structures, e.g., broken connections. We propose a novel method that learns to segment with correct topol…

cs.CV2020

Intrinsic Decomposition of Document Images In-the-Wild

Sagnik Das, Hassan Ahmed Sial, Ke Ma +3

Automatic document content processing is affected by artifacts caused by the shape of the paper, non-uniform and diverse color of lighting conditions. Fully-supervised methods on r…

cs.CV2025

CDG-MAE: Cross-view Masked Modeling using Diffusion Generated Views

Varun Belagali, Pierre Marza, Srikar Yellapragada +7

Cross-view masked autoencoding has emerged as a powerful pretext task for learning dense correspondences, which are essential for applications such as video label propagation. The…

eess.IV2019

Learning from Thresholds: Fully Automated Classification of Tumor Infiltrating Lymphocytes for Multiple Cancer Types

Shahira Abousamra, Le Hou, Rajarsi Gupta +7

Deep learning classifiers for characterization of whole slide tissue morphology require large volumes of annotated data to learn variations across different tissue and cancer types…

cs.CV2024

Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following

Qiaomu Miao, Alexandros Graikos, Jingwei Zhang +3

Training gaze following models requires a large number of images with gaze target coordinates annotated by human annotators, which is a laborious and inherently ambiguous process.…

cs.CV2023

Unsupervised and semi-supervised co-salient object detection via segmentation frequency statistics

Souradeep Chakraborty, Shujon Naha, Muhammet Bastan +2

In this paper, we address the detection of co-occurring salient objects (CoSOD) in an image group using frequency statistics in an unsupervised manner, which further enable us to d…

cs.CV2019

Weakly Labeling the Antarctic: The Penguin Colony Case

Hieu Le, Bento Gonçalves, Dimitris Samaras +1

Antarctic penguins are important ecological indicators -- especially in the face of climate change. In this work, we present a deep learning based model for semantic segmentation o…

cs.CV2026

Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention

Sounak Mondal, Dimitris Samaras, Gregory Zelinsky +1

We introduce a novel learning problem: decoding gaze into natural language descriptions of human goals across diverse visual tasks. Unlike prior work, which frames gaze decoding as…

cs.CV2025

What is the Right Embedding Space for Contrastive Learning in REC?

Kostas Triaridis, Panagiotis Kaliosis, E-Ro Nguyen +3

Referring Expression Counting (REC) requires distinguishing visually similar objects described by fine-grained text cues. Existing methods tackle this via image-text contrastive le…

cs.CV2025

PathSegDiff: Pathology Segmentation using Diffusion model representations

Sachin Kumar Danisetty, Alexandros Graikos, Srikar Yellapragada +1

Image segmentation is crucial in many computational pathology pipelines, including accurate disease diagnosis, subtyping, outcome, and survivability prediction. The common approach…

eess.IV2024

Computational Pathology: A Survey Review and The Way Forward

Mahdi S. Hosseini, Babak Ehteshami Bejnordi, Vincent Quoc-Huy Trinh +18

Computational Pathology CPath is an interdisciplinary science that augments developments of computational approaches to analyze and model medical histopathology images. The main ob…

cs.CV2020

Light Direction and Color Estimation from Single Image with Deep Regression

Hassan A. Sial, Ramon Baldrich, Maria Vanrell +1

We present a method to estimate the direction and color of the scene light source from a single image. Our method is based on two main ideas: (a) we use a new synthetic dataset wit…

cs.CV2021

Modeling Deep Learning Based Privacy Attacks on Physical Mail

Bingyao Huang, Ruyi Lian, Dimitris Samaras +1

Mail privacy protection aims to prevent unauthorized access to hidden content within an envelope since normal paper envelopes are not as safe as we think. In this paper, for the fi…

cs.CV2024

MLI-NeRF: Multi-Light Intrinsic-Aware Neural Radiance Fields

Yixiong Yang, Shilin Hu, Haoyu Wu +3

Current methods for extracting intrinsic image components, such as reflectance and shading, primarily rely on statistical priors. These methods focus mainly on simple synthetic sce…

cs.CV2026

Counting Trees from Satellite Imagery with Noisy Supervision

Dimitri Gominski, Maurice Mugabowindekwe, Qiue Xu +6

Counting individual trees is a fundamental task for environmental monitoring, yet remains largely unexplored with satellite imagery. At these resolutions, isolated trees may still…

cs.CV2020

FaceDet3D: Facial Expressions with 3D Geometric Detail Prediction

ShahRukh Athar, Albert Pumarola, Francesc Moreno-Noguer +1

Facial Expressions induce a variety of high-level details on the 3D face geometry. For example, a smile causes the wrinkling of cheeks or the formation of dimples, while being angr…

cs.CV2026

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

Minh-Quan Le, Gaurav Mittal, Cheng Zhao +3

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training…

cs.LG2019

Exascale Deep Learning to Accelerate Cancer Research

Robert M. Patton, J. Travis Johnston, Steven R. Young +9

Deep learning, through the use of neural networks, has demonstrated remarkable ability to automate many routine tasks when presented with sufficient data for training. The neural n…

cs.CV2026

Phase-Aligned RoPE for Mixed-Resolution Diffusion Transformer

Haoyu Wu, Jingyi Xu, Qiaomu Miao +2

Rotary positional embeddings (RoPE) are widely used in diffusion transformers (DiTs) to encode spatial relationships, yet their behavior with mixed-resolution tokens remains undere…

cs.CV2025

Few-shot Personalized Scanpath Prediction

Ruoyu Xue, Jingyi Xu, Sounak Mondal +4

A personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training sca…