works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
most citedChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

2 citations · 2 across the 2 of their papers we have counts for

collaborators

9 papers

cs.SE2026

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

Sizhong Qin, Yi Gu, Yao Jiang +13

The paper introduces StructureClaw, a workbench where LLM agents execute structured engineering workflows by managing typed tools and shared artifacts, and presents StructureClaw-B…

cs.CL20262 cited

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

Zhigen Li, Jianxiang Peng, Yanmeng Wang +13

Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and human-like responses, their **lack of…

cs.CL2026

TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation

Renren Jin, Tianhao Shen, Xinwei Wu +9

Conducting supervised and preference fine-tuning of large language models (LLMs) requires high-quality datasets to improve their ability to follow instructions and align with human…

cs.SE2025

OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

Tianyu Zheng, Ge Zhang, Tianhao Shen +5

The introduction of large language models has significantly advanced code generation. However, open-source models often lack the execution capabilities and iterative refinement of…

cs.AI2024

Large Language Model Safety: A Holistic Survey

Dan Shi, Tianhao Shen, Yufei Huang +10

The rapid development and deployment of large language models (LLMs) have introduced a new frontier in artificial intelligence, marked by unprecedented capabilities in natural lang…

cs.CR2024

Automated Progressive Red Teaming

Bojian Jiang, Yi Jing, Tianhao Shen +3

Ensuring the safety of large language models (LLMs) is paramount, yet identifying potential vulnerabilities is challenging. While manual red teaming is effective, it is time-consum…