works on

From the 2 of 107 linked papers with an AI index.

activity
20242026
most citedInteraction2Code: Benchmarking MLLM-based Interactive Webpage Code Generation from Interactive Prototyping

3 citations · 9 across the 31 of their papers we have counts for

collaborators
Showing cs.CLShow all

11 papers · 1 filter

cs.CL2026

Learning to Ask: When LLM Agents Meet Unclear Instruction

Wenxuan Wang, Juluan Shi, Zixuan Ling +7

Equipped with the capability to call functions, modern large language models (LLMs) can leverage external tools for addressing a range of tasks unattainable through language skills…

cs.CL2026

Data Compressibility Quantifies LLM Memorization

Yizhan Huang, Zhe Yang, Meifang Chen +3

Large Language Models (LLMs) are known to memorize portions of their training data, sometimes even reproduce content verbatim when prompted appropriately. Despite substantial inter…

cs.CL2025

ComboBench: Can LLMs Manipulate Physical Devices to Play Virtual Reality Games?

Shuqing Li, Jiayi Yan, Chenyu Niu +5

Virtual Reality (VR) games require players to translate high-level semantic actions into precise device manipulations using controllers and head-mounted displays (HMDs). While huma…

cs.CL2025

Medchain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence

Jie Liu, Wenxuan Wang, Zizhan Ma +7

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large…

cs.CL2025

Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases

Jen-tse Huang, Yuhang Yan, Linqi Liu +4

Recent failures such as Google Gemini generating people of color in Nazi-era uniforms illustrate how AI outputs can be factually plausible yet socially harmful. AI models are incre…

cs.CL2025

CLEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation

Yanyang Li, Tin Long Wong, Cheung To Hung +5

Recent advances in large language models (LLMs) have shown significant promise, yet their evaluation raises concerns, particularly regarding data contamination due to the lack of a…