activity
20242026
most citedUnderstanding Agent Scaling in LLM-Based Multi-Agent Systems via Diversity

2 citations · 3 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

LLMs Should Express Uncertainty Explicitly

Junyu Guo, Shangding Gu, Ming Jin +2

Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We study whether post-training can make a m…

cs.LG2026

StyleBench: Evaluating thinking styles in Large Language Models

Junyu Guo, Shangding Gu, Ming Jin +2

Structured reasoning can improve the inference performance of large language models (LLMs), but it also introduces computational cost and control constraints. When additional reaso…

cs.LG2026

Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization

Shangding Gu

Large language models (LLMs) are increasingly deployed in privacy-critical and personalization-oriented scenarios, yet the role of context length in shaping privacy leakage and per…

cs.LG2025

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

Shangding Gu, Xiaohan Wang, Donghao Ying +9

Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench…

cs.LG2025

Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL

Junyu Guo, Zhi Zheng, Donghao Ying +4

Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent has only a fixed dataset -- common…

cs.LG2025

Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond

Shangding Gu, Donghao Ying, Ming Jin +4

We introduce Model Feedback Learning (MFL), a novel test-time optimization framework for optimizing inputs to pre-trained AI models or deployed hardware systems without requiring a…