8 papers
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
Dawei Li, Yuguang Yao, Zhen Tan +2
Reward-guided search methods have demonstrated strong potential in enhancing tool-using agents by effectively guiding sampling and exploration over complex action spaces. As a core…
Who's Your Judge? On the Detectability of LLM-Generated Judgments
Dawei Li, Zhen Tan, Chengshuai Zhao +6
Large Language Model (LLM)-based judgments leverage powerful LLMs to efficiently evaluate candidate content and provide judgment scores. However, the inherent biases and vulnerabil…
Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era
Dawei Li, Yue Huang, Ming Li +3
Generative models such as Large Language Models, Diffusion Models, and generative adversarial networks have recently revolutionized the creation of synthetic data, offering scalabl…
Building Safer Sites: A Large-Scale Multi-Level Dataset for Construction Safety Research
Zhenhui Ou, Dawei Li, Zhen Tan +3
Construction safety research is a critical field in civil engineering, aiming to mitigate risks and prevent injuries through the analysis of site conditions and human factors. Howe…
Are Today's LLMs Ready to Explain Well-Being Concepts?
Bohan Jiang, Dawei Li, Zhen Tan +2
Well-being encompasses mental, physical, and social dimensions essential to personal growth and informed life decisions. As individuals increasingly consult Large Language Models (…
LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment
Lingyao Li, Dawei Li, Zhenhui Ou +5
Efficient simulation is essential for enhancing proactive preparedness for sudden-onset disasters such as earthquakes. Recent advancements in large language models (LLMs) as world…