activity
20242026
collaborators

9 papers

cs.AI2026

Numbers Beat Words: A Rigorous On-Premise Benchmark for Coupled MIMO Controller Tuning

Jiaxuan Chen, Haonan Li, Yang Shu

Tuning controllers for strongly coupled multi-input multi-output (MIMO) processes is difficult because decentralized auto-tuning ignores loop interaction and local optimization is…

cs.CL2026

From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves

Haritz Puerto, Haonan Li, Xudong Han +2

Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult to control and frequently violate explicit…

cs.LG2026

Training and Benchmarking Code Generation for Physics-Inspired Animations

Yanan Wang, Renxi Wang, Yongxin Wang +5

Large language models (LLMs) have been widely studied in areas such as mathematical reasoning, complex coding, and scientific problem solving. However, their ability to generate ex…

cs.CL2026

SCALAR: Scientific Citation-based Live Assessment of Long-context Academic Reasoning

Renxi Wang, Honglin Mu, Liqun Ma +5

Long-context understanding has emerged as a critical capability for large language models (LLMs). However, evaluating this ability remains challenging. We present SCALAR, a benchma…

cs.CL2025

Control Illusion: The Failure of Instruction Hierarchies in Large Language Models

Yilin Geng, Haonan Li, Honglin Mu +5

Large language models (LLMs) are increasingly deployed with hierarchical instruction schemes, where certain instructions (e.g., system-level directives) are expected to take preced…

cs.AI2025

AgentFly: Extensible and Scalable Reinforcement Learning for LM Agents

Renxi Wang, Rifo Ahmad Genadi, Bilal El Bouardi +5

Language model (LM) agents have gained significant attention for their ability to autonomously complete tasks through interactions with environments, tools, and APIs. LM agents are…