Publications (22)
GeCo-SRT: Geometry-aware Continual Adaptation for Robotic Cross-Task Sim-to-Real Transfer
Wenbo Yu, Wenke Xia, Weitao Zhang +1
Bridging the sim-to-real gap is important for applying low-cost simulation data to real-world robotic systems. However, previous methods are severely limited by treating each trans…
Enhancing Gradient Inversion Attacks in Federated Learning via Hierarchical Feature Optimization
Hao Fang, Wenbo Yu, Bin Chen +4
Federated Learning (FL) has emerged as a compelling paradigm for privacy-preserving distributed machine learning, allowing multiple clients to collaboratively train a global model…
A Probabilistic Approach to Wildfire Spread Prediction Using a Denoising Diffusion Surrogate Model
Wenbo Yu, Anirbit Ghosh, Tobias Sebastian Finn +3
Thanks to recent advances in generative AI, computers can now simulate realistic and complex natural processes. We apply this capability to predict how wildfires spread, a task mad…
Features of a nano-twist phase in the nanolayered Ti3AlC2 MAX phase
Julien Guénolé, Vincent Taupin, Maxime Vallet +2
Complex intermetallic materials known as MAX phases exhibit exceptional properties from both metals and ceramics, largely thanks to their nanolayered structure. With high-resolutio…
Every Step Evolves: Scaling Reinforcement Learning for Trillion-Scale Thinking Model
Ling Team, Anqi Shen, Baihui Li +101
We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 b…
Privacy Leakage on DNNs: A Survey of Model Inversion Attacks and Defenses
Hao Fang, Yixiang Qiu, Hongyao Yu +7
Deep Neural Networks (DNNs) have revolutionized various domains with their exceptional performance across numerous applications. However, Model Inversion (MI) attacks, which disclo…
Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation
Litao Liu, Yifan Han, Pengfei Yi +9
Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dependent and often one-to-many…
GLM-5: from Vibe Coding to Agentic Engineering
GLM-5-Team, :, Aohan Zeng +184
We present GLM-5, a next-generation foundation model designed to transition the paradigm of vibe coding to agentic engineering. Building upon the agentic, reasoning, and coding (AR…
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models
Hao Fang, Jiawei Kong, Wenbo Yu +5
Vision-Language Pre-training (VLP) models have exhibited unprecedented capability in many applications by taking full advantage of the multimodal alignment. However, previous studi…
UDA: Unsupervised Debiasing Alignment for Pair-wise LLM-as-a-Judge
Yang Zhang, Cunxiang Wang, Lindong Wu +4
Pairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own.…
Bypassing Copyright Protection in Diffusion-based Customization via Two-Stage Latent Feature Optimization
Ziang Xu, Wenbo Yu, Hongyao Yu +6
With the growing concerns over copyright infringement in diffusion-based customization, adversarial attacks have emerged as a prominent defense strategy to prevent malicious conten…
Editable-DeepSC: Cross-Modal Editable Semantic Communication Systems
Wenbo Yu, Bin Chen, Qinshan Zhang +1
Different from data-oriented communication systems that primarily focus on how to accurately transmit every bit of data, task-oriented semantic communication systems only transmit…
CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment
Wenbo Yu, Bohua Wang, Hao Fang +10
Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safety guardrails excel in English or multilin…
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
Yanzhi Tian, Cunxiang Wang, Zeming Liu +5
Large Language Models (LLMs) have significantly advanced Machine Translation (MT), applying them to linguistically complex domains-such as Social Network Services, literature etc.…
TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Traces
Shu-Xun Yang, Cunxiang Wang, Haoke Zhang +12
Agentic systems augment large language models with external tools and iterative decision making, enabling complex tasks such as deep research, function calling, and coding. However…
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
5 Team, Aohan Zeng, Xin Lv +167
We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that s…
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning
Wenke Xia, Pei Ren, Wenbo Yu +10
Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reproduction and diagnosis. Within such systems…
RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesis
Andrew Zhuoer Feng, Cunxiang Wang, Yu Luo +9
Large Language Models have evolved from single-round generators into long-horizon agents, capable of complex text synthesis scenarios. However, current evaluation frameworks lack t…
GI-NAS: Boosting Gradient Inversion Attacks Through Adaptive Neural Architecture Search
Wenbo Yu, Hao Fang, Bin Chen +5
Gradient Inversion Attacks invert the transmitted gradients in Federated Learning (FL) systems to reconstruct the sensitive data of local clients and have raised considerable priva…
MIBench: A Comprehensive Framework for Benchmarking Model Inversion Attack and Defense
Yixiang Qiu, Hongyao Yu, Hao Fang +6
Model Inversion (MI) attacks aim at leveraging the output information of target models to reconstruct privacy-sensitive training data, raising critical concerns regarding the priva…
Editable-DeepSC: Reliable Cross-Modal Semantic Communications for Facial Editing
Bin Chen, Wenbo Yu, Qinshan Zhang +4
Interactive computer vision (CV) plays a crucial role in various real-world applications, whose performance is highly dependent on communication networks. Nonetheless, the data-ori…
Mcity Data Collection for Automated Vehicles Study
Yiqun Dong, Yuanxin Zhong, Wenbo Yu +5
The main goal of this paper is to introduce the data collection effort at Mcity targeting automated vehicle development. We captured a comprehensive set of data from a set of perce…