1 citations · 2 across the 3 of their papers we have counts for
6 papers
See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
Yimeng Zhang, Jiri Gesi, Ran Xue +10
LLMs have recently demonstrated strong potential in simulating online shopper behavior. Prior work has improved action prediction by applying SFT on action traces with LLM-generate…
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
Lu Sun, Shihan Fu, Bingsheng Yao +7
Agentic AI is emerging, capable of executing tasks through natural language, such as Copilot for coding or Amazon Rufus for shopping. Evaluating these systems is challenging, as th…
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
Jiacheng Lin, Zhongruo Wang, Kun Qian +14
Supervised Fine-Tuning (SFT) on domain-specific datasets is a common approach to adapt Large Language Models (LLMs) to specialized tasks but is often believed to degrade their gene…
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
Jiaju Chen, Yuxuan Lu, Xiaojie Wang +6
Nearly all human work is collaborative; thus, the evaluation of real-world NLP applications often requires multiple dimensions that align with diverse human perspectives. As real h…
UXAgent: A System for Simulating Usability Testing of Web Design with LLM Agents
Yuxuan Lu, Bingsheng Yao, Hansu Gu +7
Usability testing is a fundamental research method that user experience (UX) researchers use to evaluate and iterate their new designs. But what about evaluating and iterating the…
UXAgent: An LLM Agent-Based Usability Testing Framework for Web Design
Yuxuan Lu, Bingsheng Yao, Hansu Gu +7
Usability testing is a fundamental yet challenging (e.g., inflexible to iterate the study design flaws and hard to recruit study participants) research method for user experience (…