papers

Publications (21)

cs.AI2026

SCMAPR: Self-Correcting Multi-Agent Prompt Refinement for Complex-Scenario Text-to-Video Generation

Chengyi Yang, Pengzhen Li, Jiayin Qi +3

Text-to-Video (T2V) generation has benefited from recent advances in diffusion models, yet current systems still struggle under complex scenarios, which are generally exacerbated b…

cs.LG2024

EffiCANet: Efficient Time Series Forecasting with Convolutional Attention

Xinxing Zhou, Jiaqi Ye, Shubao Zhao +4

The exponential growth of multivariate time series data from sensor networks in domains like industrial monitoring and smart cities requires efficient and accurate forecasting mode…

cs.CL2024

Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized Rehearsal

Jianheng Huang, Leyang Cui, Ante Wang +5

Large language models (LLMs) suffer from catastrophic forgetting during continual learning. Conventional rehearsal-based methods rely on previous training data to retain the model'…

cs.LG2024

Towards Universal Large-Scale Foundational Model for Natural Gas Demand Forecasting

Xinxing Zhou, Jiaqi Ye, Shubao Zhao +6

In the context of global energy strategy, accurate natural gas demand forecasting is crucial for ensuring efficient resource allocation and operational planning. Traditional foreca…

cs.CL2025

INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance

Shisong Chen, Qian Zhu, Wenyan Yang +15

Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capa…

cs.CL2025

Locate-and-Focus: Enhancing Terminology Translation in Speech Language Models

Suhang Wu, Jialong Tang, Chengyi Yang +6

Direct speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard,…

cs.LG2023

Efficient Training of Large-scale Industrial Fault Diagnostic Models through Federated Opportunistic Block Dropout

Yuanyuan Chen, Zichen Chen, Sheng Guo +6

Artificial intelligence (AI)-empowered industrial fault diagnostics is important in ensuring the safe operation of industrial applications. Since complex industrial systems often i…

cs.LG2023

Hierarchical Federated Learning Incentivization for Gas Usage Estimation

Has Sun, Xiaoli Tang, Chengyi Yang +5

Accurately estimating gas usage is essential for the efficient functioning of gas distribution networks and saving operational costs. Traditional methods rely on centralized data p…

cs.AI2026

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

MiniMax, :, Aili Chen +219

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The…

cs.LG2024

Wasserstein Differential Privacy

Chengyi Yang, Jiayin Qi, Aimin Zhou

Differential privacy (DP) has achieved remarkable results in the field of privacy-preserving machine learning. However, existing DP frameworks do not satisfy all the conditions for…

cs.AI2026

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

Yunbo Tang, Chengyi Yang, Shiyu Liu +4

Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these systems often suffer from a criti…

cs.CV2026

On the Robustness of Machine Unlearning for Vision-Language Models

Yujie Lin, Kaidi Jia, Jiayao Ma +2

Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work, we present the first systema…

cs.LG2023

The Prospect of Enhancing Large-Scale Heterogeneous Federated Learning with Transformers

Yulan Gao, Zhaoxiang Hou, Chengyi Yang +2

Federated learning (FL) addresses data privacy concerns by enabling collaborative training of AI models across distributed data owners. Wide adoption of FL faces the fundamental ch…

cs.LG2026

ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models

Yujie Lin, Chengyi Yang, Zhishang Xiang +2

Large language models inevitably retain sensitive information, defined as inputs that may induce harmful generations, due to training on massive web corpora, raising concerns for p…

cs.LG2023

Federated Learning in Big Model Era: Domain-Specific Multimodal Large Models

Zengxiang Li, Zhaoxiang Hou, Hui Liu +8

Multimodal data, which can comprehensively perceive and recognize the physical world, has become an essential path towards general artificial intelligence. However, multimodal larg…

cs.LG2024

HiMTM: Hierarchical Multi-Scale Masked Time Series Modeling with Self-Distillation for Long-Term Forecasting

Shubao Zhao, Ming Jin, Zhaoxiang Hou +4

Time series forecasting is a critical and challenging task in practical application. Recent advancements in pre-trained foundation models for time series forecasting have gained si…

cs.LG2026

TTCS: Test-Time Curriculum Synthesis for Self-Evolving

Chengyi Yang, Zhishang Xiang, Yunbo Tang +5

Test-Time Training offers a promising way to improve the reasoning ability of large language models (LLMs) by adapting the model using only the test questions. However, existing me…

cond-mat.mtrl-sci2025

Enhanced anomalous Hall conductivity via Ga doping in Mn\textsubscript{3}Sn and Mn\textsubscript{3}Ge

Chenyue Wen, Danrong Xiong, Chengyi Yang +2

This study examines the anomalous Hall effect (AHE) in the Heusler series \ce{Mn3Z} (Z=Ga, Ge, Sn), with a particular emphasis on the manipulation of non-collinear antiferromagneti…

cs.AR2025

STI-SNN: A 0.14 GOPS/W/PE Single-Timestep Inference FPGA-based SNN Accelerator with Algorithm and Hardware Co-Design

Kainan Wang, Chengyi Yang, Chengting Yu +3

Brain-inspired Spiking Neural Networks (SNNs) have attracted attention for their event-driven characteristics and high energy efficiency. However, the temporal dependency and irreg…

cs.CL2024

Towards Better Graph-based Cross-document Relation Extraction via Non-bridge Entity Enhancement and Prediction Debiasing

Hao Yue, Shaopeng Lai, Chengyi Yang +3

Cross-document Relation Extraction aims to predict the relation between target entities located in different documents. In this regard, the dominant models commonly retain useful i…

cs.CV2026

Object Hallucination-Free Reinforcement Unlearning for Vision-Language Models

Kaidi Jia, Yujie Lin, Chengyi Yang +2

Vision-language models (VLMs) raise growing concerns about privacy, copyright, and bias, motivating machine unlearning to remove sensitive knowledge. However, existing methods prim…