activity
20242026
most citedLADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation

2 citations · 4 across the 6 of their papers we have counts for

collaborators

9 papers

cs.CL2026

Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis

Da Song, Yuheng Huang, Boqi Chen +4

The integration of large language models (LLMs) into autonomous agents has enabled complex tool use, yet in high-stakes domains, these systems must strictly adhere to regulatory st…

cs.SE2025

TRUSTVIS: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models

Ruoyu Sun, Da Song, Jiayang Song +2

As Large Language Models (LLMs) continue to revolutionize Natural Language Processing (NLP) applications, critical concerns about their trustworthiness persist, particularly in saf…

cs.SE2025

Evaluating LLMs on Sequential API Call Through Automated Test Generation

Yuheng Huang, Jiayang Song, Da Song +4

By integrating tools from external APIs, Large Language Models (LLMs) have expanded their promising capabilities in a diverse spectrum of complex real-world tasks. However, testing…

cs.SE20252 cited

Risk Assessment Framework for Code LLMs via Leveraging Internal States

Yuheng Huang, Lei Ma, Keizaburo Nishikino +1

The pre-training paradigm plays a key role in the success of Large Language Models (LLMs), which have been recognized as one of the most significant advancements of AI recently. Bu…

cs.SE2025

Foundation Models for Autonomous Driving System: An Initial Roadmap

Xiongfei Wu, Mingfei Cheng, Xiaoning Ren +8

Recent advances in foundation models (FMs), including large language models (LLMs), vision-language models (VLMs), and world models, have opened new opportunities for autonomous dr…

cs.SE2025

Fine-grained Testing for Autonomous Driving Software: a Study on Autoware with LLM-driven Unit Testing

Wenhan Wang, Xuan Xie, Yuheng Huang +3

Testing autonomous driving systems (ADS) is critical to ensuring their reliability and safety. Existing ADS testing works focuses on designing scenarios to evaluate system-level be…