7 papers
ToolPRMBench: Evaluating and Advancing Process Reward Models for Tool-using Agents
Dawei Li, Yuguang Yao, Zhen Tan +2
Reward-guided search methods have demonstrated strong potential in enhancing tool-using agents by effectively guiding sampling and exploration over complex action spaces. As a core…
Who's Your Judge? On the Detectability of LLM-Generated Judgments
Dawei Li, Zhen Tan, Chengshuai Zhao +6
Large Language Model (LLM)-based judgments leverage powerful LLMs to efficiently evaluate candidate content and provide judgment scores. However, the inherent biases and vulnerabil…
Assessing On-the-Ground Disaster Impact Using Online Data Sources
Saketh Vishnubhatla, Ujun Jeong, Bohan Jiang +4
Assessing the impact of a disaster in terms of asset losses and human casualties is essential for preparing effective response plans. Traditional methods include offline assessment…
Building Safer Sites: A Large-Scale Multi-Level Dataset for Construction Safety Research
Zhenhui Ou, Dawei Li, Zhen Tan +3
Construction safety research is a critical field in civil engineering, aiming to mitigate risks and prevent injuries through the analysis of site conditions and human factors. Howe…
Are Today's LLMs Ready to Explain Well-Being Concepts?
Bohan Jiang, Dawei Li, Zhen Tan +2
Well-being encompasses mental, physical, and social dimensions essential to personal growth and informed life decisions. As individuals increasingly consult Large Language Models (…
Label Distribution Learning-Enhanced Dual-KNN for Text Classification
Bo Yuan, Yulin Chen, Zhen Tan +3
Many text classification methods usually introduce external information (e.g., label descriptions and knowledge bases) to improve the classification performance. Compared to extern…