activity
20182026
most citedLook Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models

48 citations · 159 across the 20 of their papers we have counts for

collaborators
Showing 2024Show all

9 papers · 1 filter

cs.SE2024★ 6 cited

Understanding the Fundamental Design Decisions of Retrieval-Augmented Generation Systems

Shengming Zhao, Yuchen Shao, Yuheng Huang +4

Retrieval-Augmented Generation (RAG) has emerged as a critical technique for enhancing large language model (LLM) capabilities. However, practitioners face significant challenges w…

cs.RO2024★ 2 cited

LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation

Zhijie Wang, Zhehua Zhou, Jiayang Song +3

Building on the advancements of Large Language Models (LLMs) and Vision Language Models (VLMs), recent research has introduced Vision-Language-Action (VLA) models as an integrated…

cs.SE2024★ 12 cited

VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation

Zhijie Wang, Zhehua Zhou, Jiayang Song +3

The rapid advancement of generative AI and multi-modal foundation models has shown significant potential in advancing robotic manipulation. Vision-language-action (VLA) models, in…

cs.SE2024

LeCov: Multi-level Testing Criteria for Large Language Models

Xuan Xie, Jiayang Song, Yuheng Huang +4

Large Language Models (LLMs) are widely used in many different domains, but because of their limited interpretability, there are questions about how trustworthy they are in various…

cs.SE2024★ 2 cited

MORTAR: A Model-based Runtime Action Repair Framework for AI-enabled Cyber-Physical Systems

Renzhi Wang, Zhehua Zhou, Jiayang Song +3

Cyber-Physical Systems (CPSs) are increasingly prevalent across various industrial and daily-life domains, with applications ranging from robotic operations to autonomous driving.…

cs.SE2024★ 1 cited

AcTracer: Active Testing of Large Language Model via Multi-Stage Sampling

Yuheng Huang, Jiayang Song, Qiang Hu +2

Performance evaluation plays a crucial role in the development life cycle of large language models (LLMs). It estimates the model's capability, elucidates behavior characteristics,…