20 papers · 1 filter
DemoShapley: Valuation of Demonstrations for In-Context Learning
Shan Xie, Man Luo, Chadly Daniel Stern +2
Large language models (LLMs) using in-context learning (ICL) excel in many tasks without task-specific fine-tuning. However, demonstration selection and ordering greatly impact ICL…
Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
Sai Shridhar Balamurali, Lu Cheng
Evaluating answers from state-of-the-art large language models (LLMs) is challenging: lexical metrics miss semantic nuances, whereas "LLM-as-Judge" scoring is computationally expen…
A Lightweight Large Language Model-Based Multi-Agent System for 2D Frame Structural Analysis
Ziheng Geng, Jiachen Liu, Ran Cao +3
Large language models (LLMs) have recently been used to empower autonomous agents in engineering, significantly improving automation and efficiency in labor-intensive workflows. Ho…
SoM-1K: A Thousand-Problem Benchmark Dataset for Strength of Materials
Qixin Wan, Zilong Wang, Jingwen Zhou +6
Foundation models have shown remarkable capabilities in various domains, but their performance on complex, multimodal engineering problems remains largely unexplored. We introduce…
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
Longchao Da, Xiaoou Liu, Jiaxin Dai +3
Understanding the uncertainty in large language model (LLM) explanations is important for evaluating their faithfulness and reasoning consistency, and thus provides insights into t…
REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models
Ke Fang, Tianyi Zhao, Lu Cheng +1
Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in safety-critical applications. Existing pr…