papers

Publications (55)

cs.CE2026

Exascale Hybrid Numerical-AI Ensembles for Operational Flood-Season Forecasting in East Asia: 15-km Decadal Hindcasts and 1-km High-Resolution Capability

Mengxuan Chen, Yunpu Xu, Qiuyan Sun +19

Seasonal forecasting of summer rainfall in East Asia remains a grand challenge, as predictability at 3 to 6 month lead times is constrained by the spring predictability barrier, we…

cs.CV2023

PKU-GoodsAD: A Supermarket Goods Dataset for Unsupervised Anomaly Detection and Segmentation

Jian Zhang, Runwei Ding, Miaoju Ban +1

Visual anomaly detection is essential and commonly used for many tasks in the field of computer vision. Recent anomaly detection datasets mainly focus on industrial automated inspe…

astro-ph.IM2023

Machine Learning for Quantum-Enhanced Gravitational-Wave Observatories

Chris Whittle, Ge Yang, Matthew Evans +1

Machine learning has become an effective tool for processing the extensive data sets produced by large physics experiments. Gravitational-wave detectors are now listening to the un…

cs.RO2022

Rapid Locomotion via Reinforcement Learning

Gabriel B Margolis, Ge Yang, Kartik Paigwar +2

Agile maneuvers such as sprinting and high-speed turning in the wild are challenging for legged robots. We present an end-to-end learned controller that achieves record agility for…

cs.AI2021

World Model as a Graph: Learning Latent Landmarks for Planning

Lunjun Zhang, Ge Yang, Bradly C. Stadie

Planning - the ability to analyze the structure of a problem in the large and decompose it into interrelated subproblems - is a hallmark of human intelligence. While deep reinforce…

cs.CL2024

LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment

Ge Yang, Changyi He, Jinyang Guo +6

Although large language models (LLMs) have demonstrated their strong intelligence ability, the high demand for computation and storage hinders their practical application. To this…

astro-ph.CO2022

Strong Lensing Source Reconstruction Using Continuous Neural Fields

Siddharth Mishra-Sharma, Ge Yang

From the nature of dark matter to the rate of expansion of our Universe, observations of distant galaxies distorted through strong gravitational lensing have the potential to answe…

cs.LG2023

Compositional Sculpting of Iterative Generative Processes

Timur Garipov, Sebastiaan De Peuter, Ge Yang +3

High training costs of generative models and the need to fine-tune them for specific tasks have created a strong interest in model reuse and composition. A key challenge in composi…

physics.bio-ph2021

Advancing biological super-resolution microscopy through deep learning: a brief review

Tianjie Yang, Yaoru Luo, Wei Ji +1

Super-resolution microscopy overcomes the diffraction limit of conventional light microscopy in spatial resolution. By providing novel spatial or spatio-temporal information on bio…

cs.RO2024

Expressive Whole-Body Control for Humanoid Robots

Xuxin Cheng, Yandong Ji, Junming Chen +3

Can we enable humanoid robots to generate rich, diverse, and expressive motions in the real world? We propose to learn a whole-body control policy on a human-sized robot to mimic h…

cs.LG2022

Overcoming the Spectral Bias of Neural Value Approximation

Ge Yang, Anurag Ajay, Pulkit Agrawal

Value approximation using deep neural networks is at the heart of off-policy deep reinforcement learning, and is often the primary module that provides learning signals to the rest…

cond-mat.mes-hall2015

Coupling an ensemble of electrons on superfluid helium to a superconducting circuit

Ge Yang, A. Fragner, G. Koolstra +4

The quantized lateral motional states and the spin states of electrons trapped on the surface of superfluid helium have been proposed as basic building blocks of a scalable quantum…

cs.CV2025

MaRS: A Fast Sampler for Mean Reverting Diffusion based on ODE and SDE Solvers

Ao Li, Wei Fang, Hongbo Zhao +3

In applications of diffusion models, controllable generation is of practical significance, but is also challenging. Current methods for controllable generation primarily focus on m…

cs.CV2019

Probabilistic Inference for Camera Calibration in Light Microscopy under Circular Motion

Yuanhao Guo, Fons J. Verbeek, Ge Yang

Robust and accurate camera calibration is essential for 3D reconstruction in light microscopy under circular motion. Conventional methods require either accurate key point matching…

cs.CV2024

XNet v2: Fewer Limitations, Better Results and Greater Universality

Yanfeng Zhou, Lingrui Li, Zichen Wang +3

XNet introduces a wavelet-based X-shaped unified architecture for fully- and semi-supervised biomedical segmentation. So far, however, XNet still faces the limitations, including p…

cs.CV2025

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Guoqing Ma, Haoyang Huang, Kun Yan +112

We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression…

cs.CV2022

Elucidating Meta-Structures of Noisy Labels in Semantic Segmentation by Deep Neural Networks

Yaoru Luo, Guole Liu, Yuanhao Guo +1

Supervised training of deep neural networks (DNNs) by noisy labels has been studied extensively in image classification but much less in image segmentation. Our understanding of th…

cs.CV2025

Control Map Distribution using Map Query Bank for Online Map Generation

Ziming Liu, Leichen Wang, Ge Yang +4

Reliable autonomous driving systems require high-definition (HD) map that contains detailed map information for planning and navigation. However, pre-build HD map requires a large…

cs.AI2023

Bilinear value networks

Zhang-Wei Hong, Ge Yang, Pulkit Agrawal

The dominant framework for off-policy multi-goal reinforcement learning involves estimating goal conditioned Q-value function. When learning to achieve multiple goals, data efficie…

cs.LG2021

Learning Task Informed Abstractions

Xiang Fu, Ge Yang, Pulkit Agrawal +1

Current model-based reinforcement learning methods struggle when operating from complex visual scenes due to their inability to prioritize task-relevant features. To mitigate this…

math.FA2019

The equivalent theorem of a new generalized Bernstein-Bezier operators

Qiu-Lan Qi, Dan-Dan Guo, Ge Yang

In this paper, a new generalized Bernstein-Bezier type operators is constructed.The estimates of the moments of these operators are investigated. The rate of convergence in terms o…

cs.CV2024

Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing

Ri-Zhao Qiu, Ge Yang, Weijia Zeng +1

Scene representations using 3D Gaussian primitives have produced excellent results in modeling the appearance of static and dynamic 3D scenes. Many graphics applications, however,…

cs.CV2022

Semi-Supervised Segmentation of Mitochondria from Electron Microscopy Images Using Spatial Continuity

Yunpeng Xiao, Youpeng Zhao, Ge Yang

Morphology of mitochondria plays critical roles in mediating their physiological functions. Accurate segmentation of mitochondria from 3D electron microscopy (EM) images is essenti…

cs.CV2023

ADFA: Attention-augmented Differentiable top-k Feature Adaptation for Unsupervised Medical Anomaly Detection

Yiming Huang, Guole Liu, Yaoru Luo +1

The scarcity of annotated data, particularly for rare diseases, limits the variability of training data and the range of detectable lesions, presenting a significant challenge for…

cs.LG2026

Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression

Yifu Ding, Jiacheng Wang, Ge Yang +4

Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead. Prior compression method…

cs.LG2023

Improving Generalization of Adversarial Training via Robust Critical Fine-Tuning

Kaijie Zhu, Jindong Wang, Xixu Hu +2

Deep neural networks are susceptible to adversarial examples, posing a significant security risk in critical applications. Adversarial Training (AT) is a well-established technique…

eess.IV2024

Representing Topological Self-Similarity Using Fractal Feature Maps for Accurate Segmentation of Tubular Structures

Jiaxing Huang, Yanfeng Zhou, Yaoru Luo +3

Accurate segmentation of long and thin tubular structures is required in a wide variety of areas such as biology, medicine, and remote sensing. The complex topology and geometry of…

cs.CV2022

Deep Neural Networks Learn Meta-Structures from Noisy Labels in Semantic Segmentation

Yaoru Luo, Guole Liu, Yuanhao Guo +1

How deep neural networks (DNNs) learn from noisy labels has been studied extensively in image classification but much less in image segmentation. So far, our understanding of the l…

cs.LG2018

Learning Plannable Representations with Causal InfoGAN

Thanard Kurutach, Aviv Tamar, Ge Yang +2

In recent years, deep generative models have been shown to 'imagine' convincing high-dimensional observations such as images, audio, and even video, learning directly from raw data…

cs.RO2025

Humanoid Policy ~ Human Policy

Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng +12

Training manipulation policies for humanoid robots with diverse data enhances their robustness and generalization across tasks and platforms. However, learning solely from robot de…

cs.RO2024

Open-TeleVision: Teleoperation with Immersive Active Visual Feedback

Xuxin Cheng, Jialong Li, Shiqi Yang +2

Teleoperation serves as a powerful method for collecting on-robot data essential for robot learning from demonstrations. The intuitiveness and ease of use of the teleoperation syst…

quant-ph2022

Single electrons on solid neon as a solid-state qubit platform

Xianjing Zhou, Gerwin Koolstra, Xufeng Zhang +9

Progress toward the realization of quantum computers requires persistent advances in their constituent building blocks - qubits. Novel qubit platforms that simultaneously embody lo…

cs.RO2026

Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation

Yajvan Ravan, Adam Rashid, Alan Yu +8

We introduce Lucid-XR, a generative data engine for creating diverse and realistic-looking multi-modal data to train real-world robotic systems. At the core of Lucid-XR is vuer, a…

cond-mat.mes-hall2019

Spectroscopy of Wigner molecules on superfluid helium using a superconducting resonator

G. Koolstra, Ge Yang, D. I. Schuster

Electrons on helium form a unique two-dimensional electron system on the interface of liquid helium and vacuum. On liquid helium, trapped electrons can arrange into strongly correl…

cs.NE2024

Toward Efficient Deep Spiking Neuron Networks:A Survey On Compression

Hui Xie, Ge Yang, Wenjuan Gao

With the rapid development of deep learning, Deep Spiking Neural Networks (DSNNs) have emerged as promising due to their unique spike event processing and asynchronous computation.…

cs.LG2022

Invariance Through Latent Alignment

Takuma Yoneda, Ge Yang, Matthew R. Walter +1

A robot's deployment environment often involves perceptual changes that differ from what it has experienced during training. Standard practices such as data augmentation attempt to…

cs.CV2026

LongCat-Next: Lexicalizing Modalities as Discrete Tokens

Meituan LongCat Team, Bin Xiao, Chao Wang +86

The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…

cs.RO2023

Neural Volumetric Memory for Visual Locomotion Control

Ruihan Yang, Ge Yang, Xiaolong Wang

Legged robots have the potential to expand the reach of autonomy beyond paved roads. In this work, we consider the difficult problem of locomotion on challenging terrains using a s…

cs.CV2023

Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation

William Shen, Ge Yang, Alan Yu +3

Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed under…

q-bio.QM2017

Deep learning based subdivision approach for large scale macromolecules structure recovery from electron cryo tomograms

Min Xu, Xiaoqi Chai, Hariank Muthakana +4

Motivation: Cellular Electron CryoTomography (CECT) enables 3D visualization of cellular organization at near-native state and in sub-molecular resolution, making it a powerful too…

physics.bio-ph2026

AI-driven Large-scale Electron Microscopy enables Whole-tissue Subcellular Digitization

Li Xiao, Liqing Liu, Hongjun Wu +7

The distribution and interactions of cellular organelles play a critical role in mediating cellular physiology and pathology. Large-scale electron microscopy enables visualization…

cs.RO2025

To the Noise and Back: Diffusion for Shared Autonomy

Takuma Yoneda, Luzhe Sun, Ge Yang +2

Shared autonomy is an operational concept in which a user and an autonomous agent collaboratively control a robotic system. It provides a number of advantages over the extremes of…

cs.LG2020

Plan2Vec: Unsupervised Representation Learning by Latent Plans

Ge Yang, Amy Zhang, Ari S. Morcos +3

In this paper we introduce plan2vec, an unsupervised representation learning approach that is inspired by reinforcement learning. Plan2vec constructs a weighted graph on an image d…

quant-ph2023

Electron charge qubits with 0.1 millisecond coherence time

Xianjing Zhou, Xinhao Li, Qianfan Chen +9

Electron charge qubits are compelling candidates for solid-state quantum computing because of their inherent simplicity in qubit design, fabrication, control, and readout. However,…

cs.RO2025

ExBody2: Advanced Expressive Humanoid Whole-Body Control

Mazeyu Ji, Xuanbin Peng, Fangchen Liu +4

This paper tackles the challenge of enabling real-world humanoid robots to perform expressive and dynamic whole-body motions while maintaining overall stability and robustness. We…

cs.RO2025

WildLMa: Long Horizon Loco-Manipulation in the Wild

Ri-Zhao Qiu, Yuchen Song, Xuanbin Peng +8

'In-the-wild' mobile manipulation aims to deploy robots in diverse real-world environments, which requires the robot to (1) have skills that generalize across object configurations…

cs.CV2025

Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos

Jiaheng Zhou, Yanfeng Zhou, Wei Fang +3

Ultrasound videos are an important form of clinical imaging data, and deep learning-based automated analysis can improve diagnostic accuracy and clinical efficiency. However, the s…

cs.AI2019

Some Considerations on Learning to Explore via Meta-Reinforcement Learning

Bradly C. Stadie, Ge Yang, Rein Houthooft +5

We consider the problem of exploration in meta reinforcement learning. Two new meta reinforcement learning algorithms are suggested: E-MAML and E-. Results are present…

cs.RO2025

Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control

Chenhao Lu, Xuxin Cheng, Jialong Li +6

Humanoid robots require both robust lower-body locomotion and precise upper-body manipulation. While recent Reinforcement Learning (RL) approaches provide whole-body loco-manipulat…

q-bio.QM2020

Few shot domain adaptation for in situ macromolecule structural classification in cryo-electron tomograms

Liangyong Yu, Ran Li, Xiangrui Zeng +5

Motivation: Cryo-Electron Tomography (cryo-ET) visualizes structure and spatial organization of macromolecules and their interactions with other subcellular components inside singl…

eess.IV2025

Opportunistic Osteoporosis Diagnosis via Texture-Preserving Self-Supervision, Mixture of Experts and Multi-Task Integration

Jiaxing Huang, Heng Guo, Le Lu +4

Osteoporosis, characterized by reduced bone mineral density (BMD) and compromised bone microstructure, increases fracture risk in aging populations. While dual-energy X-ray absorpt…

eess.IV2024

Blaze3DM: Marry Triplane Representation with Diffusion for 3D Medical Inverse Problem Solving

Jia He, Bonan Li, Ge Yang +1

Solving 3D medical inverse problems such as image restoration and reconstruction is crucial in modern medical field. However, the curse of dimensionality in 3D medical data leads m…

cs.CV2023

Robust Source-Free Domain Adaptation for Fundus Image Segmentation

Lingrui Li, Yanfeng Zhou, Ge Yang

Unsupervised Domain Adaptation (UDA) is a learning technique that transfers knowledge learned in the source domain from labelled training data to the target domain with only unlabe…

cs.RO2024

Learning Generalizable Feature Fields for Mobile Manipulation

Ri-Zhao Qiu, Yafei Hu, Yuchen Song +8

An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires c…

cs.RO2024

Learning Visual Parkour from Generated Images

Alan Yu, Ge Yang, Ran Choi +3

Fast and accurate physics simulation is an essential component of robot learning, where robots can explore failure scenarios that are difficult to produce in the real world and lea…