papers

Publications (222)

cs.RO2025

QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning

Xinyang Tong, Pengxiang Ding, Yiguo Fan +9

This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) task…

cs.LG2026

ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control

Yuzheng Hu, Ryan McKenna, Da Yu +4

Generating high-quality synthetic text under differential privacy (DP) is critical for training and evaluating language models without compromising user privacy. Prior work on synt…

cs.LG2026

Training deep physical neural networks with local physical information bottleneck

Hao Wang, Ziao Wang, Xiangpeng Liang +8

Deep learning has revolutionized modern society but faces growing energy and latency constraints. Deep physical neural networks (PNNs) are interconnected computing systems that dir…

cs.LG2025

Learning Structured Representations by Embedding Class Hierarchy with Fast Optimal Transport

Siqi Zeng, Sixian Du, Makoto Yamada +1

To embed structured knowledge within labels into feature representations, prior work [Zeng et al., 2022] proposed to use the Cophenetic Correlation Coefficient (CPCC) as a regulari…

cs.LG2021

Bridging Multi-Task Learning and Meta-Learning: Towards Efficient Training and Effective Adaptation

Haoxiang Wang, Han Zhao, Bo Li

Multi-task learning (MTL) aims to improve the generalization of several related tasks by learning them jointly. As a comparison, in addition to the joint training scheme, modern me…

eess.AS2026

MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis

Keyu An, Zhiyu Zhang, Changfeng Gao +7

This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram fram…

cs.LG2025

Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly Retraining

Weiyi Wang, Junwei Deng, Yuzheng Hu +5

Data attribution methods, which quantify the influence of individual training data points on a machine learning model, have gained increasing popularity in data-centric application…

math.OC2025

Open-source shape optimization for isogeometric shells using FEniCS and OpenMDAO

Han Zhao, John T. Hwang, Jiun-Shyan Chen

We present an open-source Python framework for the shape optimization of complex shell structures using isogeometric analysis (IGA). IGA seamlessly integrates computer-aided design…

cs.RO2026

Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process

Jiayi Chen, Wenxuan Song, Pengxiang Ding +5

Vision-language-action (VLA) models aim to understand natural language instructions and visual observations and to execute corresponding actions as an embodied agent. Recent work i…

cs.SD2026

RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS

Cong Wang, Changfeng Gao, Yang Xiang +7

Differentiable reinforcement learning (RL) frameworks like DiffRO offer a powerful approach for controllable text-to-speech (TTS), but are vulnerable to reward hacking, particularl…

cs.CL2025

AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale

Yunjie Ji, Xiaoyu Tian, Sitong Zhao +5

We present AM-Thinking-v1, a 32B dense language model that advances the frontier of reasoning, embodying the collaborative spirit of open-source innovation. Outperforming DeepSeek-…

cs.MM2022

Rethinking Task Sampling for Few-shot Vision-Language Transfer Learning

Zhenhailong Wang, Hang Yu, Manling Li +2

Despite achieving state-of-the-art zero-shot performance, existing vision-language models still fall short of few-shot transfer ability on domain-specific problems. Classical fine-…

math.NA2022

Interpolation-based immersed finite element and isogeometric analysis

Jennifer E. Fromm, Nils Wunsch, Ru Xiang +4

We introduce a new paradigm for immersed finite element and isogeometric methods based on interpolating function spaces from an unfitted background mesh into Lagrange finite elemen…

cs.AI2024

A survey of recent methods for addressing AI fairness and bias in biomedicine

Yifan Yang, Mingquan Lin, Han Zhao +3

Artificial intelligence (AI) systems have the potential to revolutionize clinical practices, including improving diagnostic accuracy and surgical decision-making, while also reduci…

cs.LG2025

Most Influential Subset Selection: Challenges, Promises, and Beyond

Yuzheng Hu, Pingbang Hu, Han Zhao +1

How can we attribute the behaviors of machine learning models to their training data? While the classic influence function sheds light on the impact of individual samples, it often…

stat.ML2024

An Empirical Study of Self-supervised Learning with Wasserstein Distance

Makoto Yamada, Yuki Takezawa, Guillaume Houry +4

In this study, we delve into the problem of self-supervised learning (SSL) utilizing the 1-Wasserstein distance on a tree structure (a.k.a., Tree-Wasserstein distance (TWD)), where…

cs.LG2017

Principled Hybrids of Generative and Discriminative Domain Adaptation

Han Zhao, Zhenyao Zhu, Junjie Hu +2

We propose a probabilistic framework for domain adaptation that blends both generative and discriminative modeling in a principled way. Under this framework, generative and discrim…

cs.LG2024

Pairwise Alignment Improves Graph Domain Adaptation

Shikun Liu, Deyu Zou, Han Zhao +1

Graph-based methods, pivotal for label inference over interconnected objects in many real-world applications, often encounter generalization challenges, if the graph used for model…

cs.LG2025

Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic

Yifei He, Yuzheng Hu, Yong Lin +2

Model merging offers an effective strategy to combine the strengths of multiple finetuned models into a unified model that preserves the specialized capabilities of each. Existing…

eess.AS2026

Qwen-Audio-3.0-Gen-Preview Technical Report

Junyu Dai, Xiaoyue Duan, Xinyue Fan +14

The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…

#audio generation#diffusion models#transformer#variational autoencoder
cs.LG2025

Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits

Weixin Chen, Han Zhao

Neural Probabilistic Circuits (NPCs), a new class of concept bottleneck models, comprise an attribute recognition model and a probabilistic circuit for reasoning. By integrating th…

cs.CV2026

LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift

Haozhe Si, Yuxuan Wan, Yuqing Wang +2

Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling, and channel dimensionality.…

cs.CV2026

CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models

Wenxuan Song, Han Zhao, Fuhao Li +7

This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard s…

cs.LG2016

A Unified Approach for Learning the Parameters of Sum-Product Networks

Han Zhao, Pascal Poupart, Geoff Gordon

We present a unified approach for learning the parameters of Sum-Product networks (SPNs). We prove that any complete and decomposable SPN is equivalent to a mixture of trees where…

cs.CL2025

Understanding Emergent In-Context Learning from a Kernel Regression Perspective

Chi Han, Ziqi Wang, Han Zhao +1

Large language models (LLMs) have initiated a paradigm shift in transfer learning. In contrast to the classic pretraining-then-finetuning procedure, in order to use LLMs for downst…

quant-ph2022

A quantum electromechanical interface for long-lived phonons

Alkim Bozkurt, Han Zhao, Chaitali Joshi +3

Controlling long-lived mechanical oscillators in the quantum regime holds promises for quantum information processing. Here, we present an electromechanical system capable of opera…

cs.LG2024

Towards Understanding the Role of Sharpness-Aware Minimization Algorithms for Out-of-Distribution Generalization

Samuel Schapiro, Han Zhao

Recently, sharpness-aware minimization (SAM) has emerged as a promising method to improve generalization by minimizing sharpness, which is known to correlate well with generalizati…

cs.LG2024

Differentially Private Post-Processing for Fair Regression

Ruicheng Xian, Qiaobo Li, Gautam Kamath +1

This paper describes a differentially private post-processing algorithm for learning fair regressors satisfying statistical parity, addressing privacy concerns of machine learning…

cs.LG2022

Greedy Modality Selection via Approximate Submodular Maximization

Runxiang Cheng, Gargi Balasubramaniam, Yifei He +2

Multimodal learning considers learning from multi-modality data, aiming to fuse heterogeneous sources of information. However, it is not always feasible to leverage all available m…

cs.LG2026

Efficient Utility-Preserving Machine Unlearning with Implicit Gradient Surgery

Shiji Zhou, Tianbai Yu, Zhi Zhang +4

Machine unlearning (MU) aims to efficiently remove sensitive or harmful memory from a pre-trained model. The key challenge is to balance the potential tradeoff between unlearning e…

cs.LG2020

Conditional Learning of Fair Representations

Han Zhao, Amanda Coston, Tameem Adel +1

We propose a novel algorithm for learning fair representations that can simultaneously mitigate two notions of disparity among different demographic subgroups in the classification…

cs.LG2024

Semi-Supervised Reward Modeling via Iterative Self-Training

Yifei He, Haoxiang Wang, Ziyan Jiang +2

Reward models (RM) capture the values and preferences of humans and play a central role in Reinforcement Learning with Human Feedback (RLHF) to align pretrained large language mode…

cs.RO2025

CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding

Wenxuan Song, Jiayi Chen, Pengxiang Ding +4

In recent years, Vision-Language-Action (VLA) models have become a vital research direction in robotics due to their impressive multimodal understanding and generalization capabili…

cs.LG2020

DynamicVAE: Decoupling Reconstruction Error and Disentangled Representation Learning

Huajie Shao, Haohong Lin, Qinmin Yang +3

This paper challenges the common assumption that the weight , in -VAE, should be larger than in order to effectively disentangle latent factors. We demonstrate that $β…

cs.RO2025

QUAR-VLA: Vision-Language-Action Model for Quadruped Robots

Pengxiang Ding, Han Zhao, Wenjie Zhang +5

The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional approaches to robot control often compartmentali…

cs.LG2025

GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection

Pingbang Hu, Joseph Melkonian, Weijing Tang +2

Gradient-based data attribution methods, such as influence functions, are critical for understanding the impact of individual training samples without requiring repeated model retr…

cs.CL2025

DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training

Xiaoyu Tian, Sitong Zhao, Haotian Wang +5

Although large language models (LLMs) have recently achieved remarkable performance on various complex reasoning benchmarks, the academic community still lacks an in-depth understa…

cs.RO2025

Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey

Shuanghao Bai, Wenxuan Song, Jiayi Chen +15

Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer vision, natural language processing, and the rise of large-scale multimodal…

cs.LG2026

Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic

Siqi Zeng, Yifei He, Meitong Liu +5

Task arithmetic, representing downstream tasks through linear operations on task vectors, has emerged as a simple yet powerful paradigm for transferring knowledge across diverse se…

eess.SY2021

Online Adaptive Optimal Control Algorithm Based on Synchronous Integral Reinforcement Learning With Explorations

Lei Guo, Han Zhao

In this paper, we present a novel algorithm named synchronous integral Q-learning, which is based on synchronous policy iteration, to solve the continuous-time infinite horizon opt…

physics.flu-dyn2026

svMultiPhysics: a finite element-based solver for cardiovascular simulations

David Codoni, Sujal Dave, David W. Parker +10

Heart disease remains the leading cause of death in the United States, motivating extensive efforts to improve its diagnosis, treatment, and prevention. Over the past decade, compu…

cs.LG2022

Algorithms and Theory for Supervised Gradual Domain Adaptation

Jing Dong, Shiji Zhou, Baoxiang Wang +1

The phenomenon of data distribution evolving over time has been observed in a range of applications, calling the needs of adaptive learning algorithms. We thus study the problem of…

cs.LG2021

Learning Invariant Representations and Risks for Semi-supervised Domain Adaptation

Bo Li, Yezhen Wang, Shanghang Zhang +4

The success of supervised learning hinges on the assumption that the training and test data come from the same underlying distribution, which is often not valid in practice due to…

cs.LG2026

Convex Dataset Valuation for Post-Training

Siqi Zeng, Christopher Jung, Rui Li +7

Improving LLM performance on downstream tasks sometimes requires leveraging auxiliary datasets during post-training. In practice, however, developers face constraints on compute, l…

cs.AI2026

CRL-VLA: Continual Vision-Language-Action Learning

Qixin Zeng, Shuo Zhang, Hongyin Zhang +6

Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Languag…

cs.LG2024

FedGTST: Boosting Global Transferability of Federated Models via Statistics Tuning

Evelyn Ma, Chao Pan, Rasoul Etesami +2

The performance of Transfer Learning (TL) heavily relies on effective pretraining, which demands large datasets and substantial computational resources. As a result, executing TL i…

cs.DC2026

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving

Xinwei Qiang, Yifan Hu, Shixuan Sun +6

Diffusion Transformers (DiTs) have become the dominant architecture for image and video generation, creating growing demand for efficient DiT serving. Existing systems assign each…

cs.OS2026

Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale

Weihao Cui, Ji Zhang, Han Zhao +5

The rapid proliferation of large language models has driven the need for efficient GPU training clusters. However, it is challenging due to the frequent occurrence of training anom…

cs.LG2025

Predicting Potential Customer Support Needs and Optimizing Search Ranking in a Two-Sided Marketplace

Do-kyum Kim, Han Zhao, Huiji Gao +3

Airbnb is an online marketplace that connects hosts and guests to unique stays and experiences. When guests stay at homes booked on Airbnb, there are a small fraction of stays that…

cs.LG2024

FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods

Xiaotian Han, Jianfeng Chi, Yu Chen +4

This paper introduces the Fair Fairness Benchmark (\textsf{FFB}), a benchmarking framework for in-processing group fairness methods. Ensuring fairness in machine learning is import…

cs.RO2024

GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot

Wenxuan Song, Han Zhao, Pengxiang Ding +4

Multi-task robot learning holds significant importance in tackling diverse and complex scenarios. However, current approaches are hindered by performance issues and difficulties in…

cs.DC2026

AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference

Di Liu, Ruitian Wang, Chen Chen +6

As large language models scale to longer contexts, loading the growing KV cache during attention computation becomes a critical bottleneck. Previous work has shown that attention c…

quant-ph2021

Strong quantum computational advantage using a superconducting quantum processor

Yulin Wu, Wan-Su Bao, Sirui Cao +51

Scaling up to a large number of qubits with high-precision control is essential in the demonstrations of quantum computational advantage to exponentially outpace the classical hard…

cs.AI2025

MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning

Jingyan Shen, Jiarui Yao, Rui Yang +5

Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, rew…

cs.CV2025

SEdit: Text-Guided Image Editing with Precise Semantic and Spatial Control

Xudong Liu, Zikun Chen, Ruowei Jiang +5

Recent advances in diffusion models have enabled high-quality generation and manipulation of images guided by texts, as well as concept learning from images. However, naive applica…

cs.RO2025

Unlock Reliable Skill Inference for Quadruped Adaptive Behavior by Skill Graph

Hongyin Zhang, Diyuan Shi, Zifeng Zhuang +6

Developing robotic intelligent systems that can adapt quickly to unseen wild situations is one of the critical challenges in pursuing autonomous robotics. Although some impressive…

cs.RO2025

VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation

Han Zhao, Jiaxuan Zhang, Wenxuan Song +2

Current vision-language-action (VLA) models, pre-trained on large-scale robotic data, exhibit strong multi-task capabilities and generalize well to variations in visual and languag…

cs.LG2022

Conditional Contrastive Learning with Kernel

Yao-Hung Hubert Tsai, Tianqin Li, Martin Q. Ma +4

Conditional contrastive learning frameworks consider the conditional sampling procedure that constructs positive or negative data pairs conditioned on specific variables. Fair cont…

cs.LG2025

MergeBench: A Benchmark for Merging Domain-Specialized LLMs

Yifei He, Siqi Zeng, Yuzheng Hu +3

Model merging provides a scalable alternative to multi-task training by combining specialized finetuned models through parameter arithmetic, enabling efficient deployment without t…

cs.OS2025

Efficient Function-as-a-Service for Large Language Models with TIDAL

Weihao Cui, Ziyi Xu, Han Zhao +4

Large Language Model (LLM) applications have emerged as a prominent use case for Function-as-a-Service (FaaS) due to their high computational demands and sporadic invocation patter…

cs.CV2025

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Han Zhao, Min Zhang, Wei Zhao +3

In recent years, the application of multimodal large language models (MLLM) in various fields has achieved remarkable success. However, as the foundation model for many downstream…

cs.LG2022

Provable Domain Generalization via Invariant-Feature Subspace Recovery

Haoxiang Wang, Haozhe Si, Bo Li +1

Domain generalization asks for models trained over a set of training environments to perform well in unseen test environments. Recently, a series of algorithms such as Invariant Ri…

cond-mat.dis-nn2020

Maximally flexible solutions of a random -satisfiability formula

Han Zhao, Hai-Jun Zhou

Random -satisfiability (-SAT) is a paradigmatic model system for studying phase transitions in constraint satisfaction problems and for developing empirical algorithms. The s…

cs.LG2024

Learning Structured Representations with Hyperbolic Embeddings

Aditya Sinha, Siqi Zeng, Makoto Yamada +1

Most real-world datasets consist of a natural hierarchy between classes or an inherent label structure that is either already available or can be constructed cheaply. However, most…

cs.CL2025

Omne-R1: Learning to Reason with Memory for Multi-hop Question Answering

Boyuan Liu, Feng Ji, Jiayan Nan +4

This paper introduces Omne-R1, a novel approach designed to enhance multi-hop question answering capabilities on schema-free knowledge graphs by integrating advanced reasoning mode…

cs.DC2024

Towards Fast Setup and High Throughput of GPU Serverless Computing

Han Zhao, Weihao Cui, Quan Chen +6

Integrating GPUs into serverless computing platforms is crucial for improving efficiency. However, existing solutions for GPU-enabled serverless computing platforms face two signif…

cs.SD2026

Eliminating stability hallucinations in llm-based tts models via attention guidance

ShiMing Wang, ZhiHao Du, Yang Xiang +6

This paper focuses on resolving stability hallucinations (e.g., repetitive or omitted speech) in LLM-based Text-to-Speech (TTS) models by improving and leveraging the attention mec…

cs.LG2022

Fundamental Limits and Tradeoffs in Invariant Representation Learning

Han Zhao, Chen Dan, Bryon Aragam +3

A wide range of machine learning applications such as privacy-preserving learning, algorithmic fairness, and domain adaptation/generalization among others, involve learning invaria…

cs.LG2026

Adaptive Online Mirror Descent for Tchebycheff Scalarization in Multi-Objective Learning

Meitong Liu, Xiaoyuan Zhang, Chulin Xie +2

Multi-objective learning (MOL) aims to learn under multiple potentially conflicting objectives and strike a proper balance. While recent preference-guided MOL methods often rely on…

cs.LG2026

Causal Neural Probabilistic Circuits

Weixin Chen, Han Zhao

Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting the class label from the concept predi…

cs.LG2024

Mitigating the Alignment Tax of RLHF

Yong Lin, Hangyu Lin, Wei Xiong +14

LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, w…

cs.LG2018

Frank-Wolfe Optimization for Symmetric-NMF under Simplicial Constraint

Han Zhao, Geoff Gordon

Symmetric nonnegative matrix factorization has found abundant applications in various domains by providing a symmetric low-rank decomposition of nonnegative matrices. In this paper…

quant-ph2024

Quantum-enabled continuous microwave-to-optics frequency conversion

Han Zhao, William David Chen, Abhishek Kejriwal +1

A quantum interface between microwave and optical photons is essential for entangling remote superconducting quantum processors. To preserve fragile quantum states, a transducer mu…

cs.RO2026

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

Huashuo Lei, Wenxuan Song, Huarui Zhang +10

Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partially observable environments.…

cs.RO2025

VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation

Wei Zhao, Pengxiang Ding, Min Zhang +4

Vision-language-action models (VLAs) have become increasingly popular in robot manipulation for their end-to-end design and remarkable performance. However, existing VLAs rely heav…

cs.RO2026

VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models

Zirui Ge, Pengxiang Ding, Baohua Yin +16

Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…

cs.CV2025

SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning

Yang Liu, Ming Ma, Xiaomin Yu +5

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for inte…

cs.AR2024

Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture

Yu Feng, Weikai Lin, Zihan Liu +6

Neural Radiance Field (NeRF) has emerged as a promising alternative for photorealistic rendering. Despite recent algorithmic advancements, achieving real-time performance on today'…

cs.LG2022

Online Continual Adaptation with Active Self-Training

Shiji Zhou, Han Zhao, Shanghang Zhang +4

Models trained with offline data often suffer from continual distribution shifts and expensive labeling in changing environments. This calls for a new online learning paradigm wher…

cs.CV2020

A Review of Single-Source Deep Unsupervised Visual Domain Adaptation

Sicheng Zhao, Xiangyu Yue, Shanghang Zhang +8

Large-scale labeled training datasets have enabled deep neural networks to excel across a wide range of benchmark vision tasks. However, in many applications, it is prohibitively e…

cs.LG2019

Efficient Multitask Feature and Relationship Learning

Han Zhao, Otilia Stretcu, Alex Smola +1

We consider a multitask learning problem, in which several predictors are learned jointly. Prior research has shown that learning the relations between tasks, and between the input…

math.PR2022

Quantitative Versions of the Two-dimensional Gaussian Product Inequalities

Ze-Chun Hu, Han Zhao, Qian-Qian Zhou

The Gaussian product inequality (GPI) conjecture is one of the most famous inequalities associated with Gaussian distributions and has attracted a lot of concerns. In this note, we…

cs.LG2023

Understanding the Impact of Adversarial Robustness on Accuracy Disparity

Yuzheng Hu, Fan Wu, Hongyang Zhang +1

While it has long been empirically observed that adversarial robustness may be at odds with standard accuracy and may have further disparate impacts on different classes, it remain…

cs.LG2021

Information Obfuscation of Graph Neural Networks

Peiyuan Liao, Han Zhao, Keyulu Xu +4

While the advent of Graph Neural Networks (GNNs) has greatly improved node and graph representation learning in many applications, the neighborhood aggregation scheme exposes addit…

cs.CL2025

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Xingxuan Li, Yao Xiao, Dianwen Ng +15

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains,…

cs.LG2021

Quantifying and Improving Transferability in Domain Generalization

Guojun Zhang, Han Zhao, Yaoliang Yu +1

Out-of-distribution generalization is one of the key challenges when transferring a model from the lab to the real world. Existing efforts mostly focus on building invariant featur…

cs.LG2020

On Learning Language-Invariant Representations for Universal Machine Translation

Han Zhao, Junjie Hu, Andrej Risteski

The goal of universal machine translation is to learn to translate between any pair of languages, given a corpus of paired translated documents for \emph{a small subset} of all pai…

cs.SD2018

Convolutional-Recurrent Neural Networks for Speech Enhancement

Han Zhao, Shuayb Zarar, Ivan Tashev +1

We propose an end-to-end model based on convolutional and recurrent neural networks for speech enhancement. Our model is purely data-driven and does not make any assumptions about…

cs.CR2022

Application of Data Encryption in Chinese Named Entity Recognition

Kaifang Long, Jikun Dong, Shengyu Fan +5

Recently, with the continuous development of deep learning, the performance of named entity recognition tasks has been dramatically improved. However, the privacy and the confident…

cs.CL2025

How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study

Yunjie Ji, Sitong Zhao, Xiaoyu Tian +5

Enhancing the reasoning capabilities of Large Language Models (LLMs) with efficiency and scalability remains a fundamental challenge in artificial intelligence research. This paper…

cs.DC2026

Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design

Chunyu Xue, Weihao Cui, Quan Chen +10

Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing…

cs.CR2025

VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes

Simon Yu, Peilin Yu, Hongbo Zheng +3

We present VISAT, a novel open dataset and benchmarking suite for evaluating model robustness in the task of traffic sign recognition with the presence of visual attributes. Built…

cs.LG2023

Train Your Own GNN Teacher: Graph-Aware Distillation on Textual Graphs

Costas Mavromatis, Vassilis N. Ioannidis, Shen Wang +6

How can we learn effective node representations on textual graphs? Graph Neural Networks (GNNs) that use Language Models (LMs) to encode textual information of graphs achieve state…

cs.LG2021

Self-supervised Representation Learning with Relative Predictive Coding

Yao-Hung Hubert Tsai, Martin Q. Ma, Muqiao Yang +3

This paper introduces Relative Predictive Coding (RPC), a new contrastive representation learning objective that maintains a good balance among training stability, minibatch size s…

physics.optics2017

Exceptional Point Engineered Glass Slide for Microscopic Thermal Mapping

Han Zhao, Zhaowei Chen, Ruogang Zhao +1

Thermal sensing with fine spatial resolution is important to the study of many scientific areas. While modern microscopy systems allow easy optical detection at high spatial resolu…

cs.IR2023

Learning List-Level Domain-Invariant Representations for Ranking

Ruicheng Xian, Honglei Zhuang, Zhen Qin +7

Domain adaptation aims to transfer the knowledge learned on (data-rich) source domains to (low-resource) target domains, and a popular method is invariant representation learning,…

cs.LG2021

Understanding and Mitigating Accuracy Disparity in Regression

Jianfeng Chi, Yuan Tian, Geoffrey J. Gordon +1

With the widespread deployment of large-scale prediction systems in high-stakes domains, e.g., face recognition, criminal justice, etc., disparity in prediction accuracy between di…

cs.CL2025

Fun-ASR Technical Report

Keyu An, Yanni Chen, Zhigao Chen +35

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep in…

cs.DC2025

Towards Resource-Efficient Serverless LLM Inference with SLINFER

Chuhao Xu, Zijun Li, Quan Chen +3

The rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow ex…

cs.RO2025

RationalVLA: A Rational Vision-Language-Action Model with Dual System

Wenxuan Song, Jiayi Chen, Wenxue Li +12

A fundamental requirement for real-world robotic deployment is the ability to understand and respond to natural language instructions. Existing language-conditioned manipulation ta…