4 papers
OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development
Li Li, Han Hu, Tianjian Zhang +27
We present OPENHARMONY BENCH, an app-level coding benchmark for evaluating LLM-based coding agents on OpenHarmony ArkTS applications. Unlike function-level benchmarks, it evaluates…
FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows
Shuo Hao, You Lu, Bihuan Chen +1
Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), tools, and control logic into…
DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents
You Lu, Kun Zhang, Bihuan Chen +1
Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critical grounding resource for LLM agents. Exi…
Lifting the Veil on Composition, Risks, and Mitigations of the Large Language Model Supply Chain
Kaifeng Huang, Bihuan Chen, You Lu +7
Large language models (LLMs) have sparked significant impact with regard to both intelligence and productivity. Numerous enterprises have integrated LLMs into their applications to…