3 papers
cs.CL2026
Shopping Reasoning Bench: An Expert-Authored Benchmark for Multi-Turn Conversational Shopping Assistants
Shuxian Fan, Seonwoo Min, Youna Hu +7
Conversational shopping assistants now serve hundreds of millions of customers, yet no existing benchmark jointly evaluates the open-ended multi-turn reasoning, domain expertise, a…
cs.AI2024
Automated Molecular Concept Generation and Labeling with Large Language Models
Zimin Zhang, Qianli Wu, Botao Xia +4
Artificial intelligence (AI) is transforming scientific research, with explainable AI methods like concept-based models (CMs) showing promise for new discoveries. However, in molec…
cs.SE2024
Does Few-Shot Learning Help LLM Performance in Code Synthesis?
Derek Xu, Tong Xie, Botao Xia +4
Large language models (LLMs) have made significant strides at code generation through improved model design, training, and chain-of-thought. However, prompt-level optimizations rem…