activity
20242026
most citedIdentifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models

3 citations · 3 across the 1 of their papers we have counts for

collaborators

8 papers

cs.SE20263 cited

Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models

Wenxuan Wang, Yuk-Kit Chan, Zixuan Ling +7

Large Language Models (LLMs) like ChatGPT are foundational in various applications due to their extensive knowledge from pre-training and fine-tuning. Despite this, they are prone…

cs.SE2025

Automatically Generating UI Code from Screenshot: A Divide-and-Conquer-Based Approach

Yuxuan Wan, Chaozheng Wang, Yi Dong +4

Websites are critical in today's digital world, with over 1.11 billion currently active and approximately 252,000 new sites launched daily. Converting website layout design into fu…

cs.AI2025

How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments

Jen-tse Huang, Eric John Li, Man Ho Lam +7

Decision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' deci…

cs.CL2024

A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

Jie Liu, Wenxuan Wang, Yihang Su +8

The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. Ho…

cs.CL2024

Understanding and Mitigating the Uncertainty in Zero-Shot Translation

Wenxuan Wang, Wenxiang Jiao, Shuo Wang +2

Zero-shot translation is a promising direction for building a comprehensive multilingual neural machine translation~(MNMT) system. However, its quality is still not satisfactory du…

cs.SE2024

LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models

Yuxuan Wan, Wenxuan Wang, Yiliu Yang +5

We introduce LogicAsker, a novel approach for evaluating and enhancing the logical reasoning capabilities of large language models (LLMs) such as ChatGPT and GPT-4. Despite LLMs' p…