7 papers · 1 filter
LLMs Should Express Uncertainty Explicitly
Junyu Guo, Shangding Gu, Ming Jin +2
Large language models (LLMs) often produce confident yet incorrect answers, which can lead to risky failures in real-world applications. We study whether post-training can make a m…
StyleBench: Evaluating thinking styles in Large Language Models
Junyu Guo, Shangding Gu, Ming Jin +2
Structured reasoning can improve the inference performance of large language models (LLMs), but it also introduces computational cost and control constraints. When additional reaso…
Representation Learning Enhanced Deep Reinforcement Learning for Optimal Operation of Hydrogen-based Multi-Energy Systems
Zhenyu Pu, Yu Yang, Lun Yang +3
Hydrogen-based multi-energy systems (HMES) have emerged as a promising low-carbon and energy-efficient solution, as it can enable the coordinated operation of electricity, heating…
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
Shangding Gu, Xiaohan Wang, Donghao Ying +9
Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench…
Don't Trade Off Safety: Diffusion Regularization for Constrained Offline RL
Junyu Guo, Zhi Zheng, Donghao Ying +4
Constrained reinforcement learning (RL) seeks high-performance policies under safety constraints. We focus on an offline setting where the agent has only a fixed dataset -- common…
Few-Shot Test-Time Optimization Without Retraining for Semiconductor Recipe Generation and Beyond
Shangding Gu, Donghao Ying, Ming Jin +4
We introduce Model Feedback Learning (MFL), a novel test-time optimization framework for optimizing inputs to pre-trained AI models or deployed hardware systems without requiring a…