Publications (55)
Exascale Hybrid Numerical-AI Ensembles for Operational Flood-Season Forecasting in East Asia: 15-km Decadal Hindcasts and 1-km High-Resolution Capability
Mengxuan Chen, Yunpu Xu, Qiuyan Sun +19
Seasonal forecasting of summer rainfall in East Asia remains a grand challenge, as predictability at 3 to 6 month lead times is constrained by the spring predictability barrier, we…
PKU-GoodsAD: A Supermarket Goods Dataset for Unsupervised Anomaly Detection and Segmentation
Jian Zhang, Runwei Ding, Miaoju Ban +1
Visual anomaly detection is essential and commonly used for many tasks in the field of computer vision. Recent anomaly detection datasets mainly focus on industrial automated inspe…
Machine Learning for Quantum-Enhanced Gravitational-Wave Observatories
Chris Whittle, Ge Yang, Matthew Evans +1
Machine learning has become an effective tool for processing the extensive data sets produced by large physics experiments. Gravitational-wave detectors are now listening to the un…
Rapid Locomotion via Reinforcement Learning
Gabriel B Margolis, Ge Yang, Kartik Paigwar +2
Agile maneuvers such as sprinting and high-speed turning in the wild are challenging for legged robots. We present an end-to-end learned controller that achieves record agility for…
World Model as a Graph: Learning Latent Landmarks for Planning
Lunjun Zhang, Ge Yang, Bradly C. Stadie
Planning - the ability to analyze the structure of a problem in the large and decompose it into interrelated subproblems - is a hallmark of human intelligence. While deep reinforce…
LLMCBench: Benchmarking Large Language Model Compression for Efficient Deployment
Ge Yang, Changyi He, Jinyang Guo +6
Although large language models (LLMs) have demonstrated their strong intelligence ability, the high demand for computation and storage hinders their practical application. To this…
Strong Lensing Source Reconstruction Using Continuous Neural Fields
Siddharth Mishra-Sharma, Ge Yang
From the nature of dark matter to the rate of expansion of our Universe, observations of distant galaxies distorted through strong gravitational lensing have the potential to answe…
Compositional Sculpting of Iterative Generative Processes
Timur Garipov, Sebastiaan De Peuter, Ge Yang +3
High training costs of generative models and the need to fine-tune them for specific tasks have created a strong interest in model reuse and composition. A key challenge in composi…
Advancing biological super-resolution microscopy through deep learning: a brief review
Tianjie Yang, Yaoru Luo, Wei Ji +1
Super-resolution microscopy overcomes the diffraction limit of conventional light microscopy in spatial resolution. By providing novel spatial or spatio-temporal information on bio…
Expressive Whole-Body Control for Humanoid Robots
Xuxin Cheng, Yandong Ji, Junming Chen +3
Can we enable humanoid robots to generate rich, diverse, and expressive motions in the real world? We propose to learn a whole-body control policy on a human-sized robot to mimic h…
Overcoming the Spectral Bias of Neural Value Approximation
Ge Yang, Anurag Ajay, Pulkit Agrawal
Value approximation using deep neural networks is at the heart of off-policy deep reinforcement learning, and is often the primary module that provides learning signals to the rest…
Coupling an ensemble of electrons on superfluid helium to a superconducting circuit
Ge Yang, A. Fragner, G. Koolstra +4
The quantized lateral motional states and the spin states of electrons trapped on the surface of superfluid helium have been proposed as basic building blocks of a scalable quantum…
MaRS: A Fast Sampler for Mean Reverting Diffusion based on ODE and SDE Solvers
Ao Li, Wei Fang, Hongbo Zhao +3
In applications of diffusion models, controllable generation is of practical significance, but is also challenging. Current methods for controllable generation primarily focus on m…
Probabilistic Inference for Camera Calibration in Light Microscopy under Circular Motion
Yuanhao Guo, Fons J. Verbeek, Ge Yang
Robust and accurate camera calibration is essential for 3D reconstruction in light microscopy under circular motion. Conventional methods require either accurate key point matching…
XNet v2: Fewer Limitations, Better Results and Greater Universality
Yanfeng Zhou, Lingrui Li, Zichen Wang +3
XNet introduces a wavelet-based X-shaped unified architecture for fully- and semi-supervised biomedical segmentation. So far, however, XNet still faces the limitations, including p…
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Guoqing Ma, Haoyang Huang, Kun Yan +112
We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression…
Elucidating Meta-Structures of Noisy Labels in Semantic Segmentation by Deep Neural Networks
Yaoru Luo, Guole Liu, Yuanhao Guo +1
Supervised training of deep neural networks (DNNs) by noisy labels has been studied extensively in image classification but much less in image segmentation. Our understanding of th…
Control Map Distribution using Map Query Bank for Online Map Generation
Ziming Liu, Leichen Wang, Ge Yang +4
Reliable autonomous driving systems require high-definition (HD) map that contains detailed map information for planning and navigation. However, pre-build HD map requires a large…
Bilinear value networks
Zhang-Wei Hong, Ge Yang, Pulkit Agrawal
The dominant framework for off-policy multi-goal reinforcement learning involves estimating goal conditioned Q-value function. When learning to achieve multiple goals, data efficie…
Learning Task Informed Abstractions
Xiang Fu, Ge Yang, Pulkit Agrawal +1
Current model-based reinforcement learning methods struggle when operating from complex visual scenes due to their inability to prioritize task-relevant features. To mitigate this…
The equivalent theorem of a new generalized Bernstein-Bezier operators
Qiu-Lan Qi, Dan-Dan Guo, Ge Yang
In this paper, a new generalized Bernstein-Bezier type operators is constructed.The estimates of the moments of these operators are investigated. The rate of convergence in terms o…
Feature Splatting: Language-Driven Physics-Based Scene Synthesis and Editing
Ri-Zhao Qiu, Ge Yang, Weijia Zeng +1
Scene representations using 3D Gaussian primitives have produced excellent results in modeling the appearance of static and dynamic 3D scenes. Many graphics applications, however,…
Semi-Supervised Segmentation of Mitochondria from Electron Microscopy Images Using Spatial Continuity
Yunpeng Xiao, Youpeng Zhao, Ge Yang
Morphology of mitochondria plays critical roles in mediating their physiological functions. Accurate segmentation of mitochondria from 3D electron microscopy (EM) images is essenti…
ADFA: Attention-augmented Differentiable top-k Feature Adaptation for Unsupervised Medical Anomaly Detection
Yiming Huang, Guole Liu, Yaoru Luo +1
The scarcity of annotated data, particularly for rare diseases, limits the variability of training data and the range of detectable lesions, presenting a significant challenge for…
Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression
Yifu Ding, Jiacheng Wang, Ge Yang +4
Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inference overhead. Prior compression method…
Improving Generalization of Adversarial Training via Robust Critical Fine-Tuning
Kaijie Zhu, Jindong Wang, Xixu Hu +2
Deep neural networks are susceptible to adversarial examples, posing a significant security risk in critical applications. Adversarial Training (AT) is a well-established technique…
Representing Topological Self-Similarity Using Fractal Feature Maps for Accurate Segmentation of Tubular Structures
Jiaxing Huang, Yanfeng Zhou, Yaoru Luo +3
Accurate segmentation of long and thin tubular structures is required in a wide variety of areas such as biology, medicine, and remote sensing. The complex topology and geometry of…
Deep Neural Networks Learn Meta-Structures from Noisy Labels in Semantic Segmentation
Yaoru Luo, Guole Liu, Yuanhao Guo +1
How deep neural networks (DNNs) learn from noisy labels has been studied extensively in image classification but much less in image segmentation. So far, our understanding of the l…
Learning Plannable Representations with Causal InfoGAN
Thanard Kurutach, Aviv Tamar, Ge Yang +2
In recent years, deep generative models have been shown to 'imagine' convincing high-dimensional observations such as images, audio, and even video, learning directly from raw data…
Humanoid Policy ~ Human Policy
Ri-Zhao Qiu, Shiqi Yang, Xuxin Cheng +12
Training manipulation policies for humanoid robots with diverse data enhances their robustness and generalization across tasks and platforms. However, learning solely from robot de…
Open-TeleVision: Teleoperation with Immersive Active Visual Feedback
Xuxin Cheng, Jialong Li, Shiqi Yang +2
Teleoperation serves as a powerful method for collecting on-robot data essential for robot learning from demonstrations. The intuitiveness and ease of use of the teleoperation syst…
Single electrons on solid neon as a solid-state qubit platform
Xianjing Zhou, Gerwin Koolstra, Xufeng Zhang +9
Progress toward the realization of quantum computers requires persistent advances in their constituent building blocks - qubits. Novel qubit platforms that simultaneously embody lo…
Lucid-XR: An Extended-Reality Data Engine for Robotic Manipulation
Yajvan Ravan, Adam Rashid, Alan Yu +8
We introduce Lucid-XR, a generative data engine for creating diverse and realistic-looking multi-modal data to train real-world robotic systems. At the core of Lucid-XR is vuer, a…
Spectroscopy of Wigner molecules on superfluid helium using a superconducting resonator
G. Koolstra, Ge Yang, D. I. Schuster
Electrons on helium form a unique two-dimensional electron system on the interface of liquid helium and vacuum. On liquid helium, trapped electrons can arrange into strongly correl…
Toward Efficient Deep Spiking Neuron Networks:A Survey On Compression
Hui Xie, Ge Yang, Wenjuan Gao
With the rapid development of deep learning, Deep Spiking Neural Networks (DSNNs) have emerged as promising due to their unique spike event processing and asynchronous computation.…
Invariance Through Latent Alignment
Takuma Yoneda, Ge Yang, Matthew R. Walter +1
A robot's deployment environment often involves perceptual changes that differ from what it has experienced during training. Standard practices such as data augmentation attempt to…
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Meituan LongCat Team, Bin Xiao, Chao Wang +86
The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…
Neural Volumetric Memory for Visual Locomotion Control
Ruihan Yang, Ge Yang, Xiaolong Wang
Legged robots have the potential to expand the reach of autonomy beyond paved roads. In this work, we consider the difficult problem of locomotion on challenging terrains using a s…
Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
William Shen, Ge Yang, Alan Yu +3
Self-supervised and language-supervised image models contain rich knowledge of the world that is important for generalization. Many robotic tasks, however, require a detailed under…
Deep learning based subdivision approach for large scale macromolecules structure recovery from electron cryo tomograms
Min Xu, Xiaoqi Chai, Hariank Muthakana +4
Motivation: Cellular Electron CryoTomography (CECT) enables 3D visualization of cellular organization at near-native state and in sub-molecular resolution, making it a powerful too…
AI-driven Large-scale Electron Microscopy enables Whole-tissue Subcellular Digitization
Li Xiao, Liqing Liu, Hongjun Wu +7
The distribution and interactions of cellular organelles play a critical role in mediating cellular physiology and pathology. Large-scale electron microscopy enables visualization…
To the Noise and Back: Diffusion for Shared Autonomy
Takuma Yoneda, Luzhe Sun, Ge Yang +2
Shared autonomy is an operational concept in which a user and an autonomous agent collaboratively control a robotic system. It provides a number of advantages over the extremes of…
Plan2Vec: Unsupervised Representation Learning by Latent Plans
Ge Yang, Amy Zhang, Ari S. Morcos +3
In this paper we introduce plan2vec, an unsupervised representation learning approach that is inspired by reinforcement learning. Plan2vec constructs a weighted graph on an image d…
Electron charge qubits with 0.1 millisecond coherence time
Xianjing Zhou, Xinhao Li, Qianfan Chen +9
Electron charge qubits are compelling candidates for solid-state quantum computing because of their inherent simplicity in qubit design, fabrication, control, and readout. However,…
ExBody2: Advanced Expressive Humanoid Whole-Body Control
Mazeyu Ji, Xuanbin Peng, Fangchen Liu +4
This paper tackles the challenge of enabling real-world humanoid robots to perform expressive and dynamic whole-body motions while maintaining overall stability and robustness. We…
WildLMa: Long Horizon Loco-Manipulation in the Wild
Ri-Zhao Qiu, Yuchen Song, Xuanbin Peng +8
'In-the-wild' mobile manipulation aims to deploy robots in diverse real-world environments, which requires the robot to (1) have skills that generalize across object configurations…
Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound Videos
Jiaheng Zhou, Yanfeng Zhou, Wei Fang +3
Ultrasound videos are an important form of clinical imaging data, and deep learning-based automated analysis can improve diagnostic accuracy and clinical efficiency. However, the s…
Some Considerations on Learning to Explore via Meta-Reinforcement Learning
Bradly C. Stadie, Ge Yang, Rein Houthooft +5
We consider the problem of exploration in meta reinforcement learning. Two new meta reinforcement learning algorithms are suggested: E-MAML and E-. Results are present…
Mobile-TeleVision: Predictive Motion Priors for Humanoid Whole-Body Control
Chenhao Lu, Xuxin Cheng, Jialong Li +6
Humanoid robots require both robust lower-body locomotion and precise upper-body manipulation. While recent Reinforcement Learning (RL) approaches provide whole-body loco-manipulat…
Few shot domain adaptation for in situ macromolecule structural classification in cryo-electron tomograms
Liangyong Yu, Ran Li, Xiangrui Zeng +5
Motivation: Cryo-Electron Tomography (cryo-ET) visualizes structure and spatial organization of macromolecules and their interactions with other subcellular components inside singl…
Opportunistic Osteoporosis Diagnosis via Texture-Preserving Self-Supervision, Mixture of Experts and Multi-Task Integration
Jiaxing Huang, Heng Guo, Le Lu +4
Osteoporosis, characterized by reduced bone mineral density (BMD) and compromised bone microstructure, increases fracture risk in aging populations. While dual-energy X-ray absorpt…
Blaze3DM: Marry Triplane Representation with Diffusion for 3D Medical Inverse Problem Solving
Jia He, Bonan Li, Ge Yang +1
Solving 3D medical inverse problems such as image restoration and reconstruction is crucial in modern medical field. However, the curse of dimensionality in 3D medical data leads m…
Robust Source-Free Domain Adaptation for Fundus Image Segmentation
Lingrui Li, Yanfeng Zhou, Ge Yang
Unsupervised Domain Adaptation (UDA) is a learning technique that transfers knowledge learned in the source domain from labelled training data to the target domain with only unlabe…
Learning Generalizable Feature Fields for Mobile Manipulation
Ri-Zhao Qiu, Yafei Hu, Yuchen Song +8
An open problem in mobile manipulation is how to represent objects and scenes in a unified manner so that robots can use both for navigation and manipulation. The latter requires c…
Learning Visual Parkour from Generated Images
Alan Yu, Ge Yang, Ran Choi +3
Fast and accurate physics simulation is an essential component of robot learning, where robots can explore failure scenarios that are difficult to produce in the real world and lea…