papers

Publications (50)

cs.CV2026

Balanced Soft mixture-of-expert model for Glaucoma Detection

Sai Venkatesh Chilukoti, Krishna Rauniyar, Min Shi +1

Glaucoma is a group of eye diseases that damage the optic nerve, often caused by elevated intraocular pressure. It is a leading cause of irreversible vision loss and is typically d…

cs.NE2025

Spiking Heterogeneous Graph Attention Networks

Buqing Cao, Qian Peng, Xiang Xie +3

Real-world graphs or networks are usually heterogeneous, involving multiple types of nodes and relationships. Heterogeneous graph neural networks (HGNNs) can effectively handle the…

cs.CV2023

When Epipolar Constraint Meets Non-local Operators in Multi-View Stereo

Tianqi Liu, Xinyi Ye, Weiyue Zhao +3

Learning-based multi-view stereo (MVS) method heavily relies on feature matching, which requires distinctive and descriptive representations. An effective solution is to apply non-…

cs.CV2022

Design What You Desire: Icon Generation from Orthogonal Application and Theme Labels

Yinpeng Chen, Zhiyu Pan, Min Shi +3

Generative adversarial networks (GANs) have been trained to be professional artists able to create stunning artworks such as face generation and image style transfer. In this paper…

cs.CV2024

NVDS+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation

Yiran Wang, Min Shi, Jiaqi Li +7

Video depth estimation aims to infer temporally consistent depth. One approach is to finetune a single-image model on each video with geometry constraints, which proves inefficient…

nucl-th2016

Probing the resonance of Dirac particle by the application of complex momentum representation

Niu Li, Min Shi, Jian-You Guo +2

Resonance plays critical roles in the formation of many physical phenomena, and several methods have been developed for the exploration of resonance. In this work, we propose a new…

cs.NE2020

Evolutionary Architecture Search for Graph Neural Networks

Min Shi, David A. Wilson, Xingquan Zhu +4

Automated machine learning (AutoML) has seen a resurgence in interest with the boom of deep learning over the past decade. In particular, Neural Architecture Search (NAS) has seen…

cs.CV2025

Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Min Shi, Fuxiao Liu, Shihao Wang +13

The ability to accurately interpret complex visual information is a crucial topic of multimodal large language models (MLLMs). Recent work indicates that enhanced visual perception…

cs.LG2019

Feature-Attention Graph Convolutional Networks for Noise Resilient Learning

Min Shi, Yufei Tang, Xingquan Zhu +1

Noise and inconsistency commonly exist in real-world information networks, due to inherent error-prone nature of human or user privacy concerns. To date, tremendous efforts have be…

cs.CV2026

Plenoptic Video Generation

Xiao Fu, Shitao Tang, Min Shi +5

Camera-controlled generative video re-rendering methods, such as ReCamMaster, have achieved remarkable progress. However, despite their success in single-view setting, these works…

cs.CV2022

Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic Counting

Min Shi, Hao Lu, Chen Feng +2

Class-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with…

cs.CV2026

Cosmos 3: Omnimodal World Models for Physical AI

NVIDIA, :, Aditi +293

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…

math.NT2026

G-companions on algebraic stacks and applications to canonical -adic local systems on Shimura stacks

Min Shi

Cases of Deligne's companion conjecture for normal schemes over finite fields have been proven by L. Lafforgue, Drinfeld, and Zheng in recent years: L. Lafforgue proved the conject…

cs.CV2026

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

Zelin Zhao, Min Shi, Bo Yuan +5

World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-Worl…

eess.IV2024

FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification

Yu Tian, Congcong Wen, Min Shi +6

Addressing fairness in artificial intelligence (AI), particularly in medical AI, is crucial for ensuring equitable healthcare outcomes. Recent efforts to enhance fairness have intr…

cs.LG2026

On Demographic Group Fairness Guarantees in Deep Learning

Yan Luo, Congcong Wen, Min Shi +3

We present a theoretical framework analyzing the relationship between data distributions and fairness guarantees in equitable deep learning. We establish novel bounds that account…

cs.LG2024

TransFair: Transferring Fairness from Ocular Disease Classification to Progression Prediction

Leila Gheisi, Henry Chu, Raju Gottumukkala +4

The use of artificial intelligence (AI) in automated disease classification significantly reduces healthcare costs and improves the accessibility of services. However, this transfo…

cs.CV2024

Geometry-aware Reconstruction and Fusion-refined Rendering for Generalizable Neural Radiance Fields

Tianqi Liu, Xinyi Ye, Min Shi +4

Generalizable NeRF aims to synthesize novel views for unseen scenes. Common practices involve constructing variance-based cost volumes for geometry reconstruction and encoding 3D d…

eess.IV2021

Fast and Accurate Single-Image Depth Estimation on Mobile Devices, Mobile AI 2021 Challenge: Report

Andrey Ignatov, Grigory Malivenko, David Plowman +35

Depth estimation is an important computer vision problem with many practical applications to mobile devices. While many solutions have been proposed for this task, they are usually…

cs.LG2021

ST-PCNN: Spatio-Temporal Physics-Coupled Neural Networks for Dynamics Forecasting

Yu Huang, James Li, Min Shi +5

Ocean current, fluid mechanics, and many other spatio-temporal physical dynamical systems are essential components of the universe. One key characteristic of such systems is that c…

cs.LG2023

EANet: Expert Attention Network for Online Trajectory Prediction

Pengfei Yao, Tianlu Mao, Min Shi +2

Trajectory prediction plays a crucial role in autonomous driving. Existing mainstream research and continuoual learning-based methods all require training on complete datasets, lea…

q-bio.GN2019

DOT: Gene-set analysis by combining decorrelated association statistics

Olga A Vsevolozhskaya, Min Shi, Fengjiao Hu +1

Historically, the majority of statistical association methods have been designed assuming availability of SNP-level information. However, modern genetic and sequencing data present…

nucl-th2018

Combination of complex momentum representation and Green's function methods in relativistic mean-field theory

Min Shi, Zhong-Ming Niu, Haozhao Liang

We have combined the complex momentum representation method with the Green's function method in the relativistic mean-field framework to establish the RMF-CMR-GF approach. This new…

nucl-th2016

Probing the resonance in the Dirac equation with quadruple-deformed potentials by complex momentum representation method

Zhi Fang, Min Shi, Jian-You Guo +3

Resonance plays critical roles in the formation of many physical phenomena, and many techniques have been developed for the exploration of resonance. In a recent letter [Phys. Rev.…

cs.GT2026

Market Games for Generative Models: Equilibria, Welfare, and Strategic Entry

Xiukun Wei, Min Shi, Xueru Zhang

Generative model ecosystems increasingly operate as competitive multi-platform markets, where platforms strategically select models from a shared pool and users with heterogeneous…

cs.CV2024

3DTopia: Large Text-to-3D Generation Model with Hybrid Diffusion Priors

Fangzhou Hong, Jiaxiang Tang, Ziang Cao +8

We present a two-stage text-to-3D generation system, namely 3DTopia, which generates high-quality general 3D assets within 5 minutes using hybrid diffusion priors. The first stage…

cs.CV2024

FairVision: Equitable Deep Learning for Eye Disease Screening via Fair Identity Scaling

Yan Luo, Muhammad Osama Khan, Yu Tian +5

Equity in AI for healthcare is crucial due to its direct impact on human well-being. Despite advancements in 2D medical imaging fairness, the fairness of 3D models remains underexp…

cs.CV2023

Harvard Glaucoma Detection and Progression: A Multimodal Multitask Dataset and Generalization-Reinforced Semi-Supervised Learning

Yan Luo, Min Shi, Yu Tian +2

Glaucoma is the number one cause of irreversible blindness globally. A major challenge for accurate glaucoma detection and progression forecasting is the bottleneck of limited labe…

cs.RO2021

Consistent Right-Invariant Fixed-Lag Smoother with Application to Visual Inertial SLAM

Jianzhu Huai, Yukai Lin, Yuan Zhuang +1

State estimation problems without absolute position measurements routinely arise in navigation of unmanned aerial vehicles, autonomous ground vehicles, etc., whose proper operation…

cs.CV2026

DuoGen: Towards General Purpose Interleaved Multimodal Generation

Min Shi, Xiaohui Zeng, Jiannan Huang +13

Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts f…

cs.CV2024

Harvard Glaucoma Fairness: A Retinal Nerve Disease Dataset for Fairness Learning and Fair Identity Normalization

Yan Luo, Yu Tian, Min Shi +5

Fairness (also known as equity interchangeably) in machine learning is important for societal well-being, but limited public datasets hinder its progress. Currently, no dedicated p…

cs.CV2022

3D Instances as 1D Kernels

Yizheng Wu, Min Shi, Shuaiyuan Du +3

We introduce a 3D instance representation, termed instance kernels, where instances are represented by one-dimensional vectors that encode the semantic, positional, and shape infor…

cs.CV2024

FairCLIP: Harnessing Fairness in Vision-Language Learning

Yan Luo, Min Shi, Muhammad Osama Khan +9

Fairness is a critical concern in deep learning, especially in healthcare, where these models influence diagnoses and treatment decisions. Although fairness has been investigated i…

cs.CV2025

Slow-Fast Architecture for Video Multi-Modal Large Language Models

Min Shi, Shihao Wang, Chieh-Yun Chen +6

Balancing temporal resolution and spatial detail under limited compute budget remains a key challenge for video-based multi-modal large language models (MLLMs). Existing methods ty…

cs.CV2026

MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models

Chieh-Yun Chen, Zhonghao Wang, Qi Chen +10

Reinforcement learning from human feedback (RLHF) with reward models has advanced alignment of generative models to human aesthetic and perceptual preferences. However, jointly opt…

cs.CV2024

FairSeg: A Large-Scale Medical Image Segmentation Dataset for Fairness Learning Using Segment Anything Model with Fair Error-Bound Scaling

Yu Tian, Min Shi, Yan Luo +3

Fairness in artificial intelligence models has gained significantly more attention in recent years, especially in the area of medicine, as fairness in medical models is critical to…

cs.CV2023

The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Weiyun Wang, Min Shi, Qingyun Li +11

We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates…

cs.CV2020

Deep Line Art Video Colorization with a Few References

Min Shi, Jia-Qi Zhang, Shu-Yu Chen +3

Coloring line art images based on the colors of reference images is an important stage in animation production, which is time-consuming and tedious. In this paper, we propose a dee…

cs.CV2025

Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light

Ali Hassani, Fengzhe Zhou, Aditya Kane +13

Many sparse attention mechanisms such as Neighborhood Attention have typically failed to consistently deliver speedup over the self attention baseline. This is largely due to the l…

cs.LG2021

Physics-Coupled Spatio-Temporal Active Learning for Dynamical Systems

Yu Huang, Yufei Tang, Xingquan Zhu +4

Spatio-temporal forecasting is of great importance in a wide range of dynamical systems applications from atmospheric science, to recent COVID-19 spread modeling. These application…

cs.CV2025

OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Shihao Wang, Zhiding Yu, Xiaohui Jiang +6

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabil…

cs.CV2022

Artifact-Tolerant Clustering-Guided Contrastive Embedding Learning for Ophthalmic Images

Min Shi, Anagha Lokhande, Mojtaba S. Fazli +10

Ophthalmic images and derivatives such as the retinal nerve fiber layer (RNFL) thickness map are crucial for detecting and monitoring ophthalmic diseases (e.g., glaucoma). For comp…

cs.SI2020

Topology and Content Co-Alignment Graph Convolutional Learning

Min Shi, Yufei Tang, Xingquan Zhu

In traditional Graph Neural Networks (GNN), graph convolutional learning is carried out through topology-driven recursive node content aggregation for network representation learni…

cs.CY2025

FairFedMed: Benchmarking Group Fairness in Federated Medical Imaging with FairLoRA

Minghan Li, Congcong Wen, Yu Tian +5

Fairness remains a critical concern in healthcare, where unequal access to services and treatment outcomes can adversely affect patient health. While Federated Learning (FL) presen…

cs.CV2025

Demystify Transformers & Convolutions in Modern Image Deep Networks

Xiaowei Hu, Min Shi, Weiyun Wang +9

Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these adva…

cs.AI2021

Deep Attributed Network Representation Learning via Attribute Enhanced Neighborhood

Cong Li, Min Shi, Bo Qu +1

Attributed network representation learning aims at learning node embeddings by integrating network structure and attribute information. It is a challenge to fully capture the micro…

cs.CV2025

OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Shihao Wang, Zhiding Yu, Xiaohui Jiang +6

The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabil…

cs.LG2019

Multi-Label Graph Convolutional Network Representation Learning

Min Shi, Yufei Tang, Xingquan Zhu +1

Knowledge representation of graph-based systems is fundamental across many disciplines. To date, most existing methods for representation learning primarily focus on networks with…

cs.CV2025

FairDiffusion: Enhancing Equity in Latent Diffusion Models via Fair Bayesian Perturbation

Yan Luo, Muhammad Osama Khan, Congcong Wen +6

Recent progress in generative AI, especially diffusion models, has demonstrated significant utility in text-to-image synthesis. Particularly in healthcare, these models offer immen…

cs.CV2025

T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation

Chieh-Yun Chen, Min Shi, Gong Zhang +1

Text-to-Image (T2I) generative models have revolutionized content creation but remain highly sensitive to prompt phrasing, often requiring users to repeatedly refine prompts multip…