collaborators

24 papers

cs.LG2026

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

Zijian Zhao, Sen Li

A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings. This architecture has been developed inde…

cs.LG2026

Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning

Zijian Zhao, Sen Li

Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically after deployment, and retraining a policy for…

cs.MA2026

Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

Zijian Zhao, Sen Li

Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many n…

cs.MA2026

RideGym: A Standardized Interface for Real-World Large-Scale Ride-Sharing System

Zijian Zhao, Yulong Hu, Sen Li

Ride-sharing has become an essential component of modern urban transportation and has attracted significant attention across computer science, transportation, and management scienc…

cs.LG2026

Pushing the Boundaries of Natural Reasoning: Interleaved Bonus from Formal-Logic Verification

Chuxue Cao, Jinluan Yang, Haoran Li +8

Large Language Models (LLMs) show remarkable capabilities, yet their stochastic next-token prediction creates logical inconsistencies and reward hacking that formal symbolic system…

cs.AI2026

Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving

Chuxue Cao, Mengze Li, Juntao Dai +7

Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas. However, their effectiveness in complex mathema…