Publications (50)
Balanced Soft mixture-of-expert model for Glaucoma Detection
Sai Venkatesh Chilukoti, Krishna Rauniyar, Min Shi +1
Glaucoma is a group of eye diseases that damage the optic nerve, often caused by elevated intraocular pressure. It is a leading cause of irreversible vision loss and is typically d…
Spiking Heterogeneous Graph Attention Networks
Buqing Cao, Qian Peng, Xiang Xie +3
Real-world graphs or networks are usually heterogeneous, involving multiple types of nodes and relationships. Heterogeneous graph neural networks (HGNNs) can effectively handle the…
When Epipolar Constraint Meets Non-local Operators in Multi-View Stereo
Tianqi Liu, Xinyi Ye, Weiyue Zhao +3
Learning-based multi-view stereo (MVS) method heavily relies on feature matching, which requires distinctive and descriptive representations. An effective solution is to apply non-…
Design What You Desire: Icon Generation from Orthogonal Application and Theme Labels
Yinpeng Chen, Zhiyu Pan, Min Shi +3
Generative adversarial networks (GANs) have been trained to be professional artists able to create stunning artworks such as face generation and image style transfer. In this paper…
NVDS+: Towards Efficient and Versatile Neural Stabilizer for Video Depth Estimation
Yiran Wang, Min Shi, Jiaqi Li +7
Video depth estimation aims to infer temporally consistent depth. One approach is to finetune a single-image model on each video with geometry constraints, which proves inefficient…
Probing the resonance of Dirac particle by the application of complex momentum representation
Niu Li, Min Shi, Jian-You Guo +2
Resonance plays critical roles in the formation of many physical phenomena, and several methods have been developed for the exploration of resonance. In this work, we propose a new…
Evolutionary Architecture Search for Graph Neural Networks
Min Shi, David A. Wilson, Xingquan Zhu +4
Automated machine learning (AutoML) has seen a resurgence in interest with the boom of deep learning over the past decade. In particular, Neural Architecture Search (NAS) has seen…
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Min Shi, Fuxiao Liu, Shihao Wang +13
The ability to accurately interpret complex visual information is a crucial topic of multimodal large language models (MLLMs). Recent work indicates that enhanced visual perception…
Feature-Attention Graph Convolutional Networks for Noise Resilient Learning
Min Shi, Yufei Tang, Xingquan Zhu +1
Noise and inconsistency commonly exist in real-world information networks, due to inherent error-prone nature of human or user privacy concerns. To date, tremendous efforts have be…
Plenoptic Video Generation
Xiao Fu, Shitao Tang, Min Shi +5
Camera-controlled generative video re-rendering methods, such as ReCamMaster, have achieved remarkable progress. However, despite their success in single-view setting, these works…
Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic Counting
Min Shi, Hao Lu, Chen Feng +2
Class-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with…
Cosmos 3: Omnimodal World Models for Physical AI
NVIDIA, :, Aditi +293
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-t…
G-companions on algebraic stacks and applications to canonical -adic local systems on Shimura stacks
Min Shi
Cases of Deligne's companion conjecture for normal schemes over finite fields have been proven by L. Lafforgue, Drinfeld, and Zheng in recent years: L. Lafforgue proved the conject…
WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling
Zelin Zhao, Min Shi, Bo Yuan +5
World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in Vision-Language-Action-Worl…
FairDomain: Achieving Fairness in Cross-Domain Medical Image Segmentation and Classification
Yu Tian, Congcong Wen, Min Shi +6
Addressing fairness in artificial intelligence (AI), particularly in medical AI, is crucial for ensuring equitable healthcare outcomes. Recent efforts to enhance fairness have intr…
On Demographic Group Fairness Guarantees in Deep Learning
Yan Luo, Congcong Wen, Min Shi +3
We present a theoretical framework analyzing the relationship between data distributions and fairness guarantees in equitable deep learning. We establish novel bounds that account…
TransFair: Transferring Fairness from Ocular Disease Classification to Progression Prediction
Leila Gheisi, Henry Chu, Raju Gottumukkala +4
The use of artificial intelligence (AI) in automated disease classification significantly reduces healthcare costs and improves the accessibility of services. However, this transfo…
Geometry-aware Reconstruction and Fusion-refined Rendering for Generalizable Neural Radiance Fields
Tianqi Liu, Xinyi Ye, Min Shi +4
Generalizable NeRF aims to synthesize novel views for unseen scenes. Common practices involve constructing variance-based cost volumes for geometry reconstruction and encoding 3D d…
Fast and Accurate Single-Image Depth Estimation on Mobile Devices, Mobile AI 2021 Challenge: Report
Andrey Ignatov, Grigory Malivenko, David Plowman +35
Depth estimation is an important computer vision problem with many practical applications to mobile devices. While many solutions have been proposed for this task, they are usually…
ST-PCNN: Spatio-Temporal Physics-Coupled Neural Networks for Dynamics Forecasting
Yu Huang, James Li, Min Shi +5
Ocean current, fluid mechanics, and many other spatio-temporal physical dynamical systems are essential components of the universe. One key characteristic of such systems is that c…
EANet: Expert Attention Network for Online Trajectory Prediction
Pengfei Yao, Tianlu Mao, Min Shi +2
Trajectory prediction plays a crucial role in autonomous driving. Existing mainstream research and continuoual learning-based methods all require training on complete datasets, lea…
DOT: Gene-set analysis by combining decorrelated association statistics
Olga A Vsevolozhskaya, Min Shi, Fengjiao Hu +1
Historically, the majority of statistical association methods have been designed assuming availability of SNP-level information. However, modern genetic and sequencing data present…
Combination of complex momentum representation and Green's function methods in relativistic mean-field theory
Min Shi, Zhong-Ming Niu, Haozhao Liang
We have combined the complex momentum representation method with the Green's function method in the relativistic mean-field framework to establish the RMF-CMR-GF approach. This new…
Probing the resonance in the Dirac equation with quadruple-deformed potentials by complex momentum representation method
Zhi Fang, Min Shi, Jian-You Guo +3
Resonance plays critical roles in the formation of many physical phenomena, and many techniques have been developed for the exploration of resonance. In a recent letter [Phys. Rev.…
Market Games for Generative Models: Equilibria, Welfare, and Strategic Entry
Xiukun Wei, Min Shi, Xueru Zhang
Generative model ecosystems increasingly operate as competitive multi-platform markets, where platforms strategically select models from a shared pool and users with heterogeneous…
3DTopia: Large Text-to-3D Generation Model with Hybrid Diffusion Priors
Fangzhou Hong, Jiaxiang Tang, Ziang Cao +8
We present a two-stage text-to-3D generation system, namely 3DTopia, which generates high-quality general 3D assets within 5 minutes using hybrid diffusion priors. The first stage…
FairVision: Equitable Deep Learning for Eye Disease Screening via Fair Identity Scaling
Yan Luo, Muhammad Osama Khan, Yu Tian +5
Equity in AI for healthcare is crucial due to its direct impact on human well-being. Despite advancements in 2D medical imaging fairness, the fairness of 3D models remains underexp…
Harvard Glaucoma Detection and Progression: A Multimodal Multitask Dataset and Generalization-Reinforced Semi-Supervised Learning
Yan Luo, Min Shi, Yu Tian +2
Glaucoma is the number one cause of irreversible blindness globally. A major challenge for accurate glaucoma detection and progression forecasting is the bottleneck of limited labe…
Consistent Right-Invariant Fixed-Lag Smoother with Application to Visual Inertial SLAM
Jianzhu Huai, Yukai Lin, Yuan Zhuang +1
State estimation problems without absolute position measurements routinely arise in navigation of unmanned aerial vehicles, autonomous ground vehicles, etc., whose proper operation…
DuoGen: Towards General Purpose Interleaved Multimodal Generation
Min Shi, Xiaohui Zeng, Jiannan Huang +13
Interleaved multimodal generation enables capabilities beyond unimodal generation models, such as step-by-step instructional guides, visual planning, and generating visual drafts f…
Harvard Glaucoma Fairness: A Retinal Nerve Disease Dataset for Fairness Learning and Fair Identity Normalization
Yan Luo, Yu Tian, Min Shi +5
Fairness (also known as equity interchangeably) in machine learning is important for societal well-being, but limited public datasets hinder its progress. Currently, no dedicated p…
3D Instances as 1D Kernels
Yizheng Wu, Min Shi, Shuaiyuan Du +3
We introduce a 3D instance representation, termed instance kernels, where instances are represented by one-dimensional vectors that encode the semantic, positional, and shape infor…
FairCLIP: Harnessing Fairness in Vision-Language Learning
Yan Luo, Min Shi, Muhammad Osama Khan +9
Fairness is a critical concern in deep learning, especially in healthcare, where these models influence diagnoses and treatment decisions. Although fairness has been investigated i…
Slow-Fast Architecture for Video Multi-Modal Large Language Models
Min Shi, Shihao Wang, Chieh-Yun Chen +6
Balancing temporal resolution and spatial detail under limited compute budget remains a key challenge for video-based multi-modal large language models (MLLMs). Existing methods ty…
MapReduce LoRA: Advancing the Pareto Front in Multi-Preference Optimization for Generative Models
Chieh-Yun Chen, Zhonghao Wang, Qi Chen +10
Reinforcement learning from human feedback (RLHF) with reward models has advanced alignment of generative models to human aesthetic and perceptual preferences. However, jointly opt…
FairSeg: A Large-Scale Medical Image Segmentation Dataset for Fairness Learning Using Segment Anything Model with Fair Error-Bound Scaling
Yu Tian, Min Shi, Yan Luo +3
Fairness in artificial intelligence models has gained significantly more attention in recent years, especially in the area of medicine, as fairness in medical models is critical to…
The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World
Weiyun Wang, Min Shi, Qingyun Li +11
We present the All-Seeing (AS) project: a large-scale data and model for recognizing and understanding everything in the open world. Using a scalable data engine that incorporates…
Deep Line Art Video Colorization with a Few References
Min Shi, Jia-Qi Zhang, Shu-Yu Chen +3
Coloring line art images based on the colors of reference images is an important stage in animation production, which is time-consuming and tedious. In this paper, we propose a dee…
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
Ali Hassani, Fengzhe Zhou, Aditya Kane +13
Many sparse attention mechanisms such as Neighborhood Attention have typically failed to consistently deliver speedup over the self attention baseline. This is largely due to the l…
Physics-Coupled Spatio-Temporal Active Learning for Dynamical Systems
Yu Huang, Yufei Tang, Xingquan Zhu +4
Spatio-temporal forecasting is of great importance in a wide range of dynamical systems applications from atmospheric science, to recent COVID-19 spread modeling. These application…
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Shihao Wang, Zhiding Yu, Xiaohui Jiang +6
The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabil…
Artifact-Tolerant Clustering-Guided Contrastive Embedding Learning for Ophthalmic Images
Min Shi, Anagha Lokhande, Mojtaba S. Fazli +10
Ophthalmic images and derivatives such as the retinal nerve fiber layer (RNFL) thickness map are crucial for detecting and monitoring ophthalmic diseases (e.g., glaucoma). For comp…
Topology and Content Co-Alignment Graph Convolutional Learning
Min Shi, Yufei Tang, Xingquan Zhu
In traditional Graph Neural Networks (GNN), graph convolutional learning is carried out through topology-driven recursive node content aggregation for network representation learni…
FairFedMed: Benchmarking Group Fairness in Federated Medical Imaging with FairLoRA
Minghan Li, Congcong Wen, Yu Tian +5
Fairness remains a critical concern in healthcare, where unequal access to services and treatment outcomes can adversely affect patient health. While Federated Learning (FL) presen…
Demystify Transformers & Convolutions in Modern Image Deep Networks
Xiaowei Hu, Min Shi, Weiyun Wang +9
Vision transformers have gained popularity recently, leading to the development of new vision backbones with improved features and consistent performance gains. However, these adva…
Deep Attributed Network Representation Learning via Attribute Enhanced Neighborhood
Cong Li, Min Shi, Bo Qu +1
Attributed network representation learning aims at learning node embeddings by integrating network structure and attribute information. It is a challenge to fully capture the micro…
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
Shihao Wang, Zhiding Yu, Xiaohui Jiang +6
The advances in vision-language models (VLMs) have led to a growing interest in autonomous driving to leverage their strong reasoning capabilities. However, extending these capabil…
Multi-Label Graph Convolutional Network Representation Learning
Min Shi, Yufei Tang, Xingquan Zhu +1
Knowledge representation of graph-based systems is fundamental across many disciplines. To date, most existing methods for representation learning primarily focus on networks with…
FairDiffusion: Enhancing Equity in Latent Diffusion Models via Fair Bayesian Perturbation
Yan Luo, Muhammad Osama Khan, Congcong Wen +6
Recent progress in generative AI, especially diffusion models, has demonstrated significant utility in text-to-image synthesis. Particularly in healthcare, these models offer immen…
T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
Chieh-Yun Chen, Min Shi, Gong Zhang +1
Text-to-Image (T2I) generative models have revolutionized content creation but remain highly sensitive to prompt phrasing, often requiring users to repeatedly refine prompts multip…