Publications (52)
Anycost GANs for Interactive Image Synthesis and Editing
Ji Lin, Richard Zhang, Frieder Ganz +2
Generative adversarial networks (GANs) have enabled photorealistic image synthesis and editing. However, due to the high computational cost of large-scale generators (e.g., StyleGA…
PoD-BIN: A Probability of Decision Bayesian Interval Design for Time-to-Event Dose-Finding Trials with Multiple Toxicity Grades
Meizi Liu, Yuan Ji, Ji Lin
We consider a Bayesian framework based on "probability of decision" for dose-finding trial designs. The proposed PoD-BIN design evaluates the posterior predictive probabilities of…
Magnetic lump motion in saturated ferromagnetic films
Xin-Wei Jin, Shi-Jie Shen, Zhan-Ying Yang +1
In this paper, we study in detail the nonlinear propagation of magnetic soliton in a ferromagnetic film. The sample is magnetized to saturation by an external field perpendicular t…
Stationary and moving bright solitons in Bose-Einstein condensates with spin-orbit coupling in a Zeeman field
JunTao He, Ji Lin
With the discovery of various matter wave solitons in spin-orbit-coupled Bose-Einstein condensates (BECs), exploring their properties has become increasingly significant. We mainly…
Rule-Guided Joint Embedding Learning over Knowledge Graphs
Qisong Li, Ji Lin, Sijia Wei +1
Recent studies on knowledge graph embedding focus on mapping entities and relations into low-dimensional vector spaces. While most existing models primarily exploit structural info…
Rogue wave, interaction solutions to the KMM system
Xin-Wei Jin, Ji Lin
In this paper, the consistent tanh expansion (CTE) method and the truncated Painlev analysis are applied to the Kraenkel-Manna-Merle (KMM) system, which describes pr…
MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
Ji Lin, Wei-Ming Chen, Han Cai +2
Tiny deep learning on microcontroller units (MCUs) is challenging due to the limited memory size. We find that the memory bottleneck is due to the imbalanced memory distribution in…
Modified Ringel-Hall algebras, naive lattice algebras and lattice algebras
Ji Lin
For a given hereditary abelian category satisfying some finiteness conditions, in certain twisted cases it is shown that the modified Ringel-Hall algebra is isomorphic to the naive…
GAN Compression: Efficient Architectures for Interactive Conditional GANs
Muyang Li, Ji Lin, Yaoyao Ding +3
Conditional Generative Adversarial Networks (cGANs) have enabled controllable image synthesis for many vision and graphics applications. However, recent cGANs are 1-2 orders of mag…
TSM: Temporal Shift Module for Efficient Video Understanding
Ji Lin, Chuang Gan, Song Han
The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computational…
Modified Ringel-Hall Algebras, Green's formula and Derived Hall Algebras
Ji Lin, Liangang Peng
In this paper we define the modified Ringel-Hall algebra $\cm\ch(\ca)$ of a hereditary abelian category $\ca$ from the category of bounded -graded co…
Construction of a new (3 + 1)-dimensional KdV equation and its closed-form solutions with solitary wave behaviour and conserved vectors
Nardjess Benoudina, Chaudry Massood Khalique, Ji Lin
This paper discusses the construction of a new -dimensional Korteweg-de Vries (KdV) equation. By employing the KdV's recursion operator, we extract two equations, and with e…
Vector rogue waves in spin-1 Bose-Einstein condensates with spin-orbit coupling
Jun-Tao He, Hui-Jun Li, Ji Lin +1
We analytically and numerically study three-component rogue waves (RWs) in spin-1 Bose-Einstein condensates with Raman-induced spin-orbit coupling (SOC). Using the multiscale pertu…
Gap solitons of the Wannier and Bloch types in spin-orbit-coupled Bose-Einstein condensates with a moiré lattice
Jun-Tao He, Xue-Ping Cheng, Xin-Wei Jin +3
Gap solitons (GSs) bifurcating from flat bands, which may be represented in terms of Wannier functions, have garnered significant interest due to their strong localization with ext…
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Guangxuan Xiao, Ji Lin, Mickael Seznec +3
Large language models (LLMs) show excellent performance but are compute- and memory-intensive. Quantization can reduce memory and accelerate inference. However, existing methods ca…
Quantum Borcherds-Bozec algebras via semi-derived Ringel-Hall algebras II: braid group actions
Ji Lin, Ming Lu, Shiquan Ruan
Based on the realization of quantum Borcherds-Bozec algebra and quantum generalized Kac-Moody algebra via semi-derived Ringel-…
TSM: Temporal Shift Module for Efficient and Scalable Video Understanding on Edge Device
Ji Lin, Chuang Gan, Kuan Wang +1
The explosive growth in video streaming requires video understanding at high accuracy and low computation cost. Conventional 2D CNNs are computationally cheap but cannot capture te…
OpenAI GPT-5 System Card
Aaditya Singh, Adam Fry, Adam Perelman +483
This is the system card published alongside the OpenAI GPT-5 launch, August 2025. GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reason…
Searching Efficient 3D Architectures with Sparse Point-Voxel Convolution
Haotian Tang, Zhijian Liu, Shengyu Zhao +4
Self-driving cars need to understand 3D scenes efficiently and accurately in order to drive safely. Given the limited hardware resources, existing 3D perception models are not able…
Lite Transformer with Long-Short Range Attention
Zhanghao Wu, Zhijian Liu, Ji Lin +2
Transformer has become ubiquitous in natural language processing (e.g., machine translation, question answering); however, it requires enormous amount of computations to achieve hi…
Reinforcement Learning from Imperfect Demonstrations
Yang Gao, Huazhe Xu, Ji Lin +3
Robust real-world learning should benefit from both demonstrations and interactions with the environment. Current approaches to learning from demonstration and reward perform super…
Tiny Machine Learning: Progress and Futures
Ji Lin, Ligeng Zhu, Wei-Ming Chen +2
Tiny Machine Learning (TinyML) is a new frontier of machine learning. By squeezing deep learning models into billions of IoT devices and microcontrollers (MCUs), we expand the scop…
Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos
Ji Lin, Chuang Gan, Song Han
Deep video recognition is more computationally expensive than image recognition, especially on large-scale datasets like Kinetics [1]. Therefore, training scalability is essential…
Design Automation for Efficient Deep Learning Computing
Song Han, Han Cai, Ligeng Zhu +4
Efficient deep learning computing requires algorithm and hardware co-design to enable specialization: we usually need to change the algorithm to reduce memory footprint and improve…
Differentiable Augmentation for Data-Efficient GAN Training
Shengyu Zhao, Zhijian Liu, Ji Lin +2
The performance of generative adversarial networks (GANs) heavily deteriorates given a limited amount of training data. This is mainly because the discriminator is memorizing the e…
A Multi-Arm Two-Stage (MATS) Design for Proof-of-Concept and Dose Optimization in Early-Phase Oncology Trials
Zhenghao Jiang, Gu Mi, Ji Lin +2
The Project Optimus initiative by the FDA's Oncology Center of Excellence is widely viewed as a groundbreaking effort to change the of conventional dose-findi…
MCUNet: Tiny Deep Learning on IoT Devices
Ji Lin, Wei-Ming Chen, Yujun Lin +3
Machine learning on tiny IoT devices based on microcontroller units (MCU) is appealing but challenging: the memory of microcontrollers is 2-3 orders of magnitude smaller even than…
Breakdown of the correspondence between the real-complex and delocalization-localization transitions in non-Hermitian quasicrystals
Wen Chen, Shujie Cheng, Ji Lin +2
The correspondence between the real-complex transition in energy and delocalization-localization transition is well-established in a class of Aubry-Andr'e-Harper model with exponen…
Alternative Analysis Methods for Time to Event Endpoints under Non-proportional Hazards: A Comparative Analysis
Ray S. Lin, Ji Lin, Satrajit Roychoudhury +16
The log-rank test is most powerful under proportional hazards (PH). In practice, non-PH patterns are often observed in clinical trials, such as in immuno-oncology; therefore, alter…
Joint Monocular 3D Vehicle Detection and Tracking
Hou-Ning Hu, Qi-Zhi Cai, Dequan Wang +5
Vehicle 3D extents and trajectories are critical cues for predicting the future location of vehicles and planning future agent ego-motion based on those predictions. In this paper,…
Defensive Quantization: When Efficiency Meets Robustness
Ji Lin, Chuang Gan, Song Han
Neural network quantization is becoming an industry standard to efficiently deploy deep learning models on hardware platforms, such as CPU, GPU, TPU, and FPGAs. However, we observe…
From Green's formula to Derived Hall algebras
Ji Lin
The aim of this note is to clarify the relationship between Green's formula and the associativity of multiplication for derived Hall algebra in the sense of Toën (Duke Math J 135(…
GPT-4o System Card
OpenAI, :, Aaron Hurst +416
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's…
VILA: On Pre-training for Visual Language Models
Ji Lin, Hongxu Yin, Wei Ping +7
Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM wi…
On-Device Training Under 256KB Memory
Ji Lin, Ligeng Zhu, Wei-Ming Chen +3
On-device training enables the model to adapt to new data collected from the sensors by fine-tuning a pre-trained model. Users can benefit from customized AI models without having…
A Seamless Phase II/III Design with Dose Optimization for Oncology Drug Development
Yuhan Li, Yiding Zhang, Gu Mi +1
The US FDA's Project Optimus initiative that emphasizes dose optimization prior to marketing approval represents a pivotal shift in oncology drug development. It has a ripple effec…
Efficient Spatially Sparse Inference for Conditional GANs and Diffusion Models
Muyang Li, Ji Lin, Chenlin Meng +3
During image editing, existing deep generative models tend to re-synthesize the entire output from scratch, including the unedited regions. This leads to a significant waste of com…
AMC: AutoML for Model Compression and Acceleration on Mobile Devices
Yihui He, Ji Lin, Zhijian Liu +3
Model compression is a critical technique to efficiently deploy neural network models on mobile devices which have limited computation resources and tight power budgets. Convention…
Semi-derived Ringel-Hall algebras and Hall algebras of odd-periodic relative derived categories
Ji Lin, Liangang Peng
Let be a positive integer and a hereditary abelian category satisfying some finiteness conditions. We define the semi-derived Ringel-Hall algebra of …
Offsite-Tuning: Transfer Learning without Full Model
Guangxuan Xiao, Ji Lin, Song Han
Transfer learning is important for foundation models to adapt to downstream tasks. However, many foundation models are proprietary, so users must share their data with model owners…
Deep Clustering based Boundary-Decoder Net for Inter and Intra Layer Stress Prediction of Heterogeneous Integrated IC Chip
Kart Leong Lim, Ji Lin
High stress occurs when 3D heterogeneous IC packages are subjected to thermal cycling at extreme temperatures. Stress mainly occurs at the interface between different materials. We…
Novel Gravastar Solutions: Investigating Stability, Energy, and Entropy in the Presence of Cloud of Strings and Quintessence
Faisal Javed, Ji Lin
Gravastars, theoretical alternatives to black holes, have captured the interest of scientists in astrophysics due to their unique properties. This paper aims to further investigate…
Cyclotron dynamics of a Bose-Einstein condensate in a quadruple-well potential with synthetic gauge fields
Wen-Yuan Wang, Ji Lin, Jie Liu
We investigate the cyclotron dynamics of Bose-Einstein condensate (BEC) in a quadruple-well potential with synthetic gauge fields. We use laser-assisted tunneling to generate large…
APQ: Joint Search for Network Architecture, Pruning and Quantization Policy
Tianzhe Wang, Kuan Wang, Han Cai +3
We present APQ for efficient deep learning inference on resource-constrained hardware. Unlike previous methods that separately search the neural architecture, pruning policy, and q…
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang +7
Large language models (LLMs) have transformed numerous AI applications. On-device LLM is becoming increasingly important: running LLMs locally on edge devices can reduce the cloud…
PockEngine: Sparse and Efficient Fine-tuning in a Pocket
Ligeng Zhu, Lanxiang Hu, Ji Lin +4
On-device learning and efficient fine-tuning enable continuous and privacy-preserving customization (e.g., locally fine-tuning large language models on personalized data). However,…
Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
Han Cai, Ji Lin, Yujun Lin +5
Deep neural networks (DNNs) have achieved unprecedented success in the field of artificial intelligence (AI), including computer vision, natural language processing and speech reco…
Network Augmentation for Tiny Deep Learning
Han Cai, Chuang Gan, Ji Lin +1
We introduce Network Augmentation (NetAug), a new training method for improving the performance of tiny neural networks. Existing regularization techniques (e.g., data augmentation…
Integration of Efficacy Biomarkers Together with Toxicity Endpoints in Immune-Oncology Dose Finding Studies
Yiding Zhang, Zhixing Xu, Hui Quan +1
The primary objective of phase I oncology studies is to establish the safety profile of a new treatment and determine the maximum tolerated dose (MTD). This is motivated by the dev…
Hardware-Centric AutoML for Mixed-Precision Quantization
Kuan Wang, Zhijian Liu, Yujun Lin +2
Model quantization is a widely used technique to compress and accelerate deep neural network (DNN) inference. Emergent DNN hardware accelerators begin to support mixed precision (1…
HAQ: Hardware-Aware Automated Quantization with Mixed Precision
Kuan Wang, Zhijian Liu, Yujun Lin +2
Model quantization is a widely used technique to compress and accelerate deep neural network (DNN) inference. Emergent DNN hardware accelerators begin to support mixed precision (1…
gpt-oss-120b & gpt-oss-20b Model Card
OpenAI, :, Sandhini Agarwal +124
We present gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models that push the frontier of accuracy and inference cost. The models use an efficient mixture-of-expert trans…