Publications (222)
QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
Xinyang Tong, Pengxiang Ding, Yiguo Fan +9
This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) task…
ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control
Yuzheng Hu, Ryan McKenna, Da Yu +4
Generating high-quality synthetic text under differential privacy (DP) is critical for training and evaluating language models without compromising user privacy. Prior work on synt…
Training deep physical neural networks with local physical information bottleneck
Hao Wang, Ziao Wang, Xiangpeng Liang +8
Deep learning has revolutionized modern society but faces growing energy and latency constraints. Deep physical neural networks (PNNs) are interconnected computing systems that dir…
Learning Structured Representations by Embedding Class Hierarchy with Fast Optimal Transport
Siqi Zeng, Sixian Du, Makoto Yamada +1
To embed structured knowledge within labels into feature representations, prior work [Zeng et al., 2022] proposed to use the Cophenetic Correlation Coefficient (CPCC) as a regulari…
Bridging Multi-Task Learning and Meta-Learning: Towards Efficient Training and Effective Adaptation
Haoxiang Wang, Han Zhao, Bo Li
Multi-task learning (MTL) aims to improve the generalization of several related tasks by learning them jointly. As a comparison, in addition to the joint training scheme, modern me…
MELA-TTS: Joint transformer-diffusion model with representation alignment for speech synthesis
Keyu An, Zhiyu Zhang, Changfeng Gao +7
This work introduces MELA-TTS, a novel joint transformer-diffusion framework for end-to-end text-to-speech synthesis. By autoregressively generating continuous mel-spectrogram fram…
Taming Hyperparameter Sensitivity in Data Attribution: Practical Selection Without Costly Retraining
Weiyi Wang, Junwei Deng, Yuzheng Hu +5
Data attribution methods, which quantify the influence of individual training data points on a machine learning model, have gained increasing popularity in data-centric application…
Open-source shape optimization for isogeometric shells using FEniCS and OpenMDAO
Han Zhao, John T. Hwang, Jiun-Shyan Chen
We present an open-source Python framework for the shape optimization of complex shell structures using isogeometric analysis (IGA). IGA seamlessly integrates computer-aided design…
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
Jiayi Chen, Wenxuan Song, Pengxiang Ding +5
Vision-language-action (VLA) models aim to understand natural language instructions and visual observations and to execute corresponding actions as an embodied agent. Recent work i…
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
Cong Wang, Changfeng Gao, Yang Xiang +7
Differentiable reinforcement learning (RL) frameworks like DiffRO offer a powerful approach for controllable text-to-speech (TTS), but are vulnerable to reward hacking, particularl…
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale
Yunjie Ji, Xiaoyu Tian, Sitong Zhao +5
We present AM-Thinking-v1, a 32B dense language model that advances the frontier of reasoning, embodying the collaborative spirit of open-source innovation. Outperforming DeepSeek-…
Rethinking Task Sampling for Few-shot Vision-Language Transfer Learning
Zhenhailong Wang, Hang Yu, Manling Li +2
Despite achieving state-of-the-art zero-shot performance, existing vision-language models still fall short of few-shot transfer ability on domain-specific problems. Classical fine-…
Interpolation-based immersed finite element and isogeometric analysis
Jennifer E. Fromm, Nils Wunsch, Ru Xiang +4
We introduce a new paradigm for immersed finite element and isogeometric methods based on interpolating function spaces from an unfitted background mesh into Lagrange finite elemen…
A survey of recent methods for addressing AI fairness and bias in biomedicine
Yifan Yang, Mingquan Lin, Han Zhao +3
Artificial intelligence (AI) systems have the potential to revolutionize clinical practices, including improving diagnostic accuracy and surgical decision-making, while also reduci…
Most Influential Subset Selection: Challenges, Promises, and Beyond
Yuzheng Hu, Pingbang Hu, Han Zhao +1
How can we attribute the behaviors of machine learning models to their training data? While the classic influence function sheds light on the impact of individual samples, it often…
An Empirical Study of Self-supervised Learning with Wasserstein Distance
Makoto Yamada, Yuki Takezawa, Guillaume Houry +4
In this study, we delve into the problem of self-supervised learning (SSL) utilizing the 1-Wasserstein distance on a tree structure (a.k.a., Tree-Wasserstein distance (TWD)), where…
Principled Hybrids of Generative and Discriminative Domain Adaptation
Han Zhao, Zhenyao Zhu, Junjie Hu +2
We propose a probabilistic framework for domain adaptation that blends both generative and discriminative modeling in a principled way. Under this framework, generative and discrim…
Pairwise Alignment Improves Graph Domain Adaptation
Shikun Liu, Deyu Zou, Han Zhao +1
Graph-based methods, pivotal for label inference over interconnected objects in many real-world applications, often encounter generalization challenges, if the graph used for model…
Localize-and-Stitch: Efficient Model Merging via Sparse Task Arithmetic
Yifei He, Yuzheng Hu, Yong Lin +2
Model merging offers an effective strategy to combine the strengths of multiple finetuned models into a unified model that preserves the specialized capabilities of each. Existing…
Qwen-Audio-3.0-Gen-Preview Technical Report
Junyu Dai, Xiaoyue Duan, Xinyue Fan +14
The paper introduces Qwen-Audio-3.0-Gen-Preview, a unified non‑autoregressive model that uses a diffusion transformer and a shared VAE to generate complete mixed‑waveform audio fro…
Understanding and Improving Adversarial Robustness of Neural Probabilistic Circuits
Weixin Chen, Han Zhao
Neural Probabilistic Circuits (NPCs), a new class of concept bottleneck models, comprise an attribute recognition model and a probabilistic circuit for reasoning. By integrating th…
LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift
Haozhe Si, Yuxuan Wan, Yuqing Wang +2
Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling, and channel dimensionality.…
CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Models
Wenxuan Song, Han Zhao, Fuhao Li +7
This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard s…
A Unified Approach for Learning the Parameters of Sum-Product Networks
Han Zhao, Pascal Poupart, Geoff Gordon
We present a unified approach for learning the parameters of Sum-Product networks (SPNs). We prove that any complete and decomposable SPN is equivalent to a mixture of trees where…
Understanding Emergent In-Context Learning from a Kernel Regression Perspective
Chi Han, Ziqi Wang, Han Zhao +1
Large language models (LLMs) have initiated a paradigm shift in transfer learning. In contrast to the classic pretraining-then-finetuning procedure, in order to use LLMs for downst…
A quantum electromechanical interface for long-lived phonons
Alkim Bozkurt, Han Zhao, Chaitali Joshi +3
Controlling long-lived mechanical oscillators in the quantum regime holds promises for quantum information processing. Here, we present an electromechanical system capable of opera…
Towards Understanding the Role of Sharpness-Aware Minimization Algorithms for Out-of-Distribution Generalization
Samuel Schapiro, Han Zhao
Recently, sharpness-aware minimization (SAM) has emerged as a promising method to improve generalization by minimizing sharpness, which is known to correlate well with generalizati…
Differentially Private Post-Processing for Fair Regression
Ruicheng Xian, Qiaobo Li, Gautam Kamath +1
This paper describes a differentially private post-processing algorithm for learning fair regressors satisfying statistical parity, addressing privacy concerns of machine learning…
Greedy Modality Selection via Approximate Submodular Maximization
Runxiang Cheng, Gargi Balasubramaniam, Yifei He +2
Multimodal learning considers learning from multi-modality data, aiming to fuse heterogeneous sources of information. However, it is not always feasible to leverage all available m…
Efficient Utility-Preserving Machine Unlearning with Implicit Gradient Surgery
Shiji Zhou, Tianbai Yu, Zhi Zhang +4
Machine unlearning (MU) aims to efficiently remove sensitive or harmful memory from a pre-trained model. The key challenge is to balance the potential tradeoff between unlearning e…
Conditional Learning of Fair Representations
Han Zhao, Amanda Coston, Tameem Adel +1
We propose a novel algorithm for learning fair representations that can simultaneously mitigate two notions of disparity among different demographic subgroups in the classification…
Semi-Supervised Reward Modeling via Iterative Self-Training
Yifei He, Haoxiang Wang, Ziyan Jiang +2
Reward models (RM) capture the values and preferences of humans and play a central role in Reinforcement Learning with Human Feedback (RLHF) to align pretrained large language mode…
CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
Wenxuan Song, Jiayi Chen, Pengxiang Ding +4
In recent years, Vision-Language-Action (VLA) models have become a vital research direction in robotics due to their impressive multimodal understanding and generalization capabili…
DynamicVAE: Decoupling Reconstruction Error and Disentangled Representation Learning
Huajie Shao, Haohong Lin, Qinmin Yang +3
This paper challenges the common assumption that the weight , in -VAE, should be larger than in order to effectively disentangle latent factors. We demonstrate that $β…
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
Pengxiang Ding, Han Zhao, Wenjie Zhang +5
The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional approaches to robot control often compartmentali…
GraSS: Scalable Data Attribution with Gradient Sparsification and Sparse Projection
Pingbang Hu, Joseph Melkonian, Weijing Tang +2
Gradient-based data attribution methods, such as influence functions, are critical for understanding the impact of individual training samples without requiring repeated model retr…
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training
Xiaoyu Tian, Sitong Zhao, Haotian Wang +5
Although large language models (LLMs) have recently achieved remarkable performance on various complex reasoning benchmarks, the academic community still lacks an in-depth understa…
Towards a Unified Understanding of Robot Manipulation: A Comprehensive Survey
Shuanghao Bai, Wenxuan Song, Jiayi Chen +15
Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer vision, natural language processing, and the rise of large-scale multimodal…
Task Vector Bases: A Unified and Scalable Framework for Compressed Task Arithmetic
Siqi Zeng, Yifei He, Meitong Liu +5
Task arithmetic, representing downstream tasks through linear operations on task vectors, has emerged as a simple yet powerful paradigm for transferring knowledge across diverse se…
Online Adaptive Optimal Control Algorithm Based on Synchronous Integral Reinforcement Learning With Explorations
Lei Guo, Han Zhao
In this paper, we present a novel algorithm named synchronous integral Q-learning, which is based on synchronous policy iteration, to solve the continuous-time infinite horizon opt…
svMultiPhysics: a finite element-based solver for cardiovascular simulations
David Codoni, Sujal Dave, David W. Parker +10
Heart disease remains the leading cause of death in the United States, motivating extensive efforts to improve its diagnosis, treatment, and prevention. Over the past decade, compu…
Algorithms and Theory for Supervised Gradual Domain Adaptation
Jing Dong, Shiji Zhou, Baoxiang Wang +1
The phenomenon of data distribution evolving over time has been observed in a range of applications, calling the needs of adaptive learning algorithms. We thus study the problem of…
Learning Invariant Representations and Risks for Semi-supervised Domain Adaptation
Bo Li, Yezhen Wang, Shanghang Zhang +4
The success of supervised learning hinges on the assumption that the training and test data come from the same underlying distribution, which is often not valid in practice due to…
Convex Dataset Valuation for Post-Training
Siqi Zeng, Christopher Jung, Rui Li +7
Improving LLM performance on downstream tasks sometimes requires leveraging auxiliary datasets during post-training. In practice, however, developers face constraints on compute, l…
CRL-VLA: Continual Vision-Language-Action Learning
Qixin Zeng, Shuo Zhang, Hongyin Zhang +6
Lifelong learning is critical for embodied agents in open-world environments, where reinforcement learning fine-tuning has emerged as an important paradigm to enable Vision-Languag…
FedGTST: Boosting Global Transferability of Federated Models via Statistics Tuning
Evelyn Ma, Chao Pan, Rasoul Etesami +2
The performance of Transfer Learning (TL) heavily relies on effective pretraining, which demands large datasets and substantial computational resources. As a result, executing TL i…
GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving
Xinwei Qiang, Yifan Hu, Shixuan Sun +6
Diffusion Transformers (DiTs) have become the dominant architecture for image and video generation, creating growing demand for efficient DiT serving. Existing systems assign each…
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
Weihao Cui, Ji Zhang, Han Zhao +5
The rapid proliferation of large language models has driven the need for efficient GPU training clusters. However, it is challenging due to the frequent occurrence of training anom…
Predicting Potential Customer Support Needs and Optimizing Search Ranking in a Two-Sided Marketplace
Do-kyum Kim, Han Zhao, Huiji Gao +3
Airbnb is an online marketplace that connects hosts and guests to unique stays and experiences. When guests stay at homes booked on Airbnb, there are a small fraction of stays that…
FFB: A Fair Fairness Benchmark for In-Processing Group Fairness Methods
Xiaotian Han, Jianfeng Chi, Yu Chen +4
This paper introduces the Fair Fairness Benchmark (\textsf{FFB}), a benchmarking framework for in-processing group fairness methods. Ensuring fairness in machine learning is import…
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
Wenxuan Song, Han Zhao, Pengxiang Ding +4
Multi-task robot learning holds significant importance in tackling diverse and complex scenarios. However, current approaches are hindered by performance issues and difficulties in…
AB-Sparse: Sparse Attention with Adaptive Block Size for Accurate and Efficient Long-Context Inference
Di Liu, Ruitian Wang, Chen Chen +6
As large language models scale to longer contexts, loading the growing KV cache during attention computation becomes a critical bottleneck. Previous work has shown that attention c…
Strong quantum computational advantage using a superconducting quantum processor
Yulin Wu, Wan-Su Bao, Sirui Cao +51
Scaling up to a large number of qubits with high-precision control is essential in the demonstrations of quantum computational advantage to exponentially outpace the classical hard…
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
Jingyan Shen, Jiarui Yao, Rui Yang +5
Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, rew…
SEdit: Text-Guided Image Editing with Precise Semantic and Spatial Control
Xudong Liu, Zikun Chen, Ruowei Jiang +5
Recent advances in diffusion models have enabled high-quality generation and manipulation of images guided by texts, as well as concept learning from images. However, naive applica…
Unlock Reliable Skill Inference for Quadruped Adaptive Behavior by Skill Graph
Hongyin Zhang, Diyuan Shi, Zifeng Zhuang +6
Developing robotic intelligent systems that can adapt quickly to unseen wild situations is one of the critical challenges in pursuing autonomous robotics. Although some impressive…
VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
Han Zhao, Jiaxuan Zhang, Wenxuan Song +2
Current vision-language-action (VLA) models, pre-trained on large-scale robotic data, exhibit strong multi-task capabilities and generalize well to variations in visual and languag…
Conditional Contrastive Learning with Kernel
Yao-Hung Hubert Tsai, Tianqin Li, Martin Q. Ma +4
Conditional contrastive learning frameworks consider the conditional sampling procedure that constructs positive or negative data pairs conditioned on specific variables. Fair cont…
MergeBench: A Benchmark for Merging Domain-Specialized LLMs
Yifei He, Siqi Zeng, Yuzheng Hu +3
Model merging provides a scalable alternative to multi-task training by combining specialized finetuned models through parameter arithmetic, enabling efficient deployment without t…
Efficient Function-as-a-Service for Large Language Models with TIDAL
Weihao Cui, Ziyi Xu, Han Zhao +4
Large Language Model (LLM) applications have emerged as a prominent use case for Function-as-a-Service (FaaS) due to their high computational demands and sporadic invocation patter…
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
Han Zhao, Min Zhang, Wei Zhao +3
In recent years, the application of multimodal large language models (MLLM) in various fields has achieved remarkable success. However, as the foundation model for many downstream…
Provable Domain Generalization via Invariant-Feature Subspace Recovery
Haoxiang Wang, Haozhe Si, Bo Li +1
Domain generalization asks for models trained over a set of training environments to perform well in unseen test environments. Recently, a series of algorithms such as Invariant Ri…
Maximally flexible solutions of a random -satisfiability formula
Han Zhao, Hai-Jun Zhou
Random -satisfiability (-SAT) is a paradigmatic model system for studying phase transitions in constraint satisfaction problems and for developing empirical algorithms. The s…
Learning Structured Representations with Hyperbolic Embeddings
Aditya Sinha, Siqi Zeng, Makoto Yamada +1
Most real-world datasets consist of a natural hierarchy between classes or an inherent label structure that is either already available or can be constructed cheaply. However, most…
Omne-R1: Learning to Reason with Memory for Multi-hop Question Answering
Boyuan Liu, Feng Ji, Jiayan Nan +4
This paper introduces Omne-R1, a novel approach designed to enhance multi-hop question answering capabilities on schema-free knowledge graphs by integrating advanced reasoning mode…
Towards Fast Setup and High Throughput of GPU Serverless Computing
Han Zhao, Weihao Cui, Quan Chen +6
Integrating GPUs into serverless computing platforms is crucial for improving efficiency. However, existing solutions for GPU-enabled serverless computing platforms face two signif…
Eliminating stability hallucinations in llm-based tts models via attention guidance
ShiMing Wang, ZhiHao Du, Yang Xiang +6
This paper focuses on resolving stability hallucinations (e.g., repetitive or omitted speech) in LLM-based Text-to-Speech (TTS) models by improving and leveraging the attention mec…
Fundamental Limits and Tradeoffs in Invariant Representation Learning
Han Zhao, Chen Dan, Bryon Aragam +3
A wide range of machine learning applications such as privacy-preserving learning, algorithmic fairness, and domain adaptation/generalization among others, involve learning invaria…
Adaptive Online Mirror Descent for Tchebycheff Scalarization in Multi-Objective Learning
Meitong Liu, Xiaoyuan Zhang, Chulin Xie +2
Multi-objective learning (MOL) aims to learn under multiple potentially conflicting objectives and strike a proper balance. While recent preference-guided MOL methods often rely on…
Causal Neural Probabilistic Circuits
Weixin Chen, Han Zhao
Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting the class label from the concept predi…
Mitigating the Alignment Tax of RLHF
Yong Lin, Hangyu Lin, Wei Xiong +14
LLMs acquire a wide range of abilities during pre-training, but aligning LLMs under Reinforcement Learning with Human Feedback (RLHF) can lead to forgetting pretrained abilities, w…
Frank-Wolfe Optimization for Symmetric-NMF under Simplicial Constraint
Han Zhao, Geoff Gordon
Symmetric nonnegative matrix factorization has found abundant applications in various domains by providing a symmetric low-rank decomposition of nonnegative matrices. In this paper…
Quantum-enabled continuous microwave-to-optics frequency conversion
Han Zhao, William David Chen, Abhishek Kejriwal +1
A quantum interface between microwave and optical photons is essential for entangling remote superconducting quantum processors. To preserve fragile quantum states, a transducer mu…
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
Huashuo Lei, Wenxuan Song, Huarui Zhang +10
Memory is a critical component of robotic intelligence, as robots must rely on past observations and actions to accomplish long-horizon tasks in partially observable environments.…
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
Wei Zhao, Pengxiang Ding, Min Zhang +4
Vision-language-action models (VLAs) have become increasingly popular in robot manipulation for their end-to-end design and remarkable performance. However, existing VLAs rely heav…
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
Zirui Ge, Pengxiang Ding, Baohua Yin +16
Video action models are an appealing foundation for Vision--Language--Action systems because they can learn visual dynamics from large-scale video data and transfer this knowledge…
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
Yang Liu, Ming Ma, Xiaomin Yu +5
Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for inte…
Potamoi: Accelerating Neural Rendering via a Unified Streaming Architecture
Yu Feng, Weikai Lin, Zihan Liu +6
Neural Radiance Field (NeRF) has emerged as a promising alternative for photorealistic rendering. Despite recent algorithmic advancements, achieving real-time performance on today'…
Online Continual Adaptation with Active Self-Training
Shiji Zhou, Han Zhao, Shanghang Zhang +4
Models trained with offline data often suffer from continual distribution shifts and expensive labeling in changing environments. This calls for a new online learning paradigm wher…
A Review of Single-Source Deep Unsupervised Visual Domain Adaptation
Sicheng Zhao, Xiangyu Yue, Shanghang Zhang +8
Large-scale labeled training datasets have enabled deep neural networks to excel across a wide range of benchmark vision tasks. However, in many applications, it is prohibitively e…
Efficient Multitask Feature and Relationship Learning
Han Zhao, Otilia Stretcu, Alex Smola +1
We consider a multitask learning problem, in which several predictors are learned jointly. Prior research has shown that learning the relations between tasks, and between the input…
Quantitative Versions of the Two-dimensional Gaussian Product Inequalities
Ze-Chun Hu, Han Zhao, Qian-Qian Zhou
The Gaussian product inequality (GPI) conjecture is one of the most famous inequalities associated with Gaussian distributions and has attracted a lot of concerns. In this note, we…
Understanding the Impact of Adversarial Robustness on Accuracy Disparity
Yuzheng Hu, Fan Wu, Hongyang Zhang +1
While it has long been empirically observed that adversarial robustness may be at odds with standard accuracy and may have further disparate impacts on different classes, it remain…
Information Obfuscation of Graph Neural Networks
Peiyuan Liao, Han Zhao, Keyulu Xu +4
While the advent of Graph Neural Networks (GNNs) has greatly improved node and graph representation learning in many applications, the neighborhood aggregation scheme exposes addit…
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
Xingxuan Li, Yao Xiao, Dianwen Ng +15
Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains,…
Quantifying and Improving Transferability in Domain Generalization
Guojun Zhang, Han Zhao, Yaoliang Yu +1
Out-of-distribution generalization is one of the key challenges when transferring a model from the lab to the real world. Existing efforts mostly focus on building invariant featur…
On Learning Language-Invariant Representations for Universal Machine Translation
Han Zhao, Junjie Hu, Andrej Risteski
The goal of universal machine translation is to learn to translate between any pair of languages, given a corpus of paired translated documents for \emph{a small subset} of all pai…
Convolutional-Recurrent Neural Networks for Speech Enhancement
Han Zhao, Shuayb Zarar, Ivan Tashev +1
We propose an end-to-end model based on convolutional and recurrent neural networks for speech enhancement. Our model is purely data-driven and does not make any assumptions about…
Application of Data Encryption in Chinese Named Entity Recognition
Kaifang Long, Jikun Dong, Shengyu Fan +5
Recently, with the continuous development of deep learning, the performance of named entity recognition tasks has been dramatically improved. However, the privacy and the confident…
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
Yunjie Ji, Sitong Zhao, Xiaoyu Tian +5
Enhancing the reasoning capabilities of Large Language Models (LLMs) with efficiency and scalability remains a fundamental challenge in artificial intelligence research. This paper…
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
Chunyu Xue, Weihao Cui, Quan Chen +10
Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing…
VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes
Simon Yu, Peilin Yu, Hongbo Zheng +3
We present VISAT, a novel open dataset and benchmarking suite for evaluating model robustness in the task of traffic sign recognition with the presence of visual attributes. Built…
Train Your Own GNN Teacher: Graph-Aware Distillation on Textual Graphs
Costas Mavromatis, Vassilis N. Ioannidis, Shen Wang +6
How can we learn effective node representations on textual graphs? Graph Neural Networks (GNNs) that use Language Models (LMs) to encode textual information of graphs achieve state…
Self-supervised Representation Learning with Relative Predictive Coding
Yao-Hung Hubert Tsai, Martin Q. Ma, Muqiao Yang +3
This paper introduces Relative Predictive Coding (RPC), a new contrastive representation learning objective that maintains a good balance among training stability, minibatch size s…
Exceptional Point Engineered Glass Slide for Microscopic Thermal Mapping
Han Zhao, Zhaowei Chen, Ruogang Zhao +1
Thermal sensing with fine spatial resolution is important to the study of many scientific areas. While modern microscopy systems allow easy optical detection at high spatial resolu…
Learning List-Level Domain-Invariant Representations for Ranking
Ruicheng Xian, Honglei Zhuang, Zhen Qin +7
Domain adaptation aims to transfer the knowledge learned on (data-rich) source domains to (low-resource) target domains, and a popular method is invariant representation learning,…
Understanding and Mitigating Accuracy Disparity in Regression
Jianfeng Chi, Yuan Tian, Geoffrey J. Gordon +1
With the widespread deployment of large-scale prediction systems in high-stakes domains, e.g., face recognition, criminal justice, etc., disparity in prediction accuracy between di…
Fun-ASR Technical Report
Keyu An, Yanni Chen, Zhigao Chen +35
In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model size scaling, and deep in…
Towards Resource-Efficient Serverless LLM Inference with SLINFER
Chuhao Xu, Zijun Li, Quan Chen +3
The rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow ex…
RationalVLA: A Rational Vision-Language-Action Model with Dual System
Wenxuan Song, Jiayi Chen, Wenxue Li +12
A fundamental requirement for real-world robotic deployment is the ability to understand and respond to natural language instructions. Existing language-conditioned manipulation ta…