activity
20242026
collaborators

8 papers

cs.LG2026

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning

Zhao Yang, Yuxuan Jiang, Ting-Chih Chen +18

Reinforcement learning (RL) has become central to LLM post-training, yet the methods that dominate current pipelines, PPO and GRPO, represent only a narrow slice of what RL offers.…

cs.LG2026

Diffusion and Flow Matching Models for Tabular Data: A Survey

Zhong Li, Qi Huang, Lincen Yang +5

Deep generative models have made rapid progress in image, text, audio, and video generation, and are increasingly being applied to structured records. For tabular data, however, ge…

cs.AI2026

MM-OptBench: A Solver-Grounded Benchmark for Multimodal Optimization Modeling

Zhong Li, Qi Huang, Yuxuan Zhu +6

Optimization modeling translates real decision-making problems into mathematical optimization models and solver-executable implementations. Although language models are increasingl…

cs.LG2026

From Heuristic Selection to Automated Algorithm Design: LLMs Benefit from Strong Priors

Qi Huang, Furong Ye, Ananta Shahane +2

Large Language Models (LLMs) have already been widely adopted for automated algorithm design, demonstrating strong abilities in generating and evolving algorithms across various fi…

cs.LG2025

Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching

Zhong Li, Qi Huang, Yuxuan Zhu +4

We introduce Time-Conditioned Contraction Matching (TCCM), a novel method for semi-supervised anomaly detection in tabular data. TCCM is inspired by flow matching, a recent generat…

cs.AI2025

Why Are You Wrong? Counterfactual Explanations for Language Grounding with 3D Objects

Tobias Preintner, Weixuan Yuan, Qi Huang +4

Combining natural language and geometric shapes is an emerging research area with multiple applications in robotics and language-assisted design. A crucial task in this domain is o…