5 papers
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Qiming Shi, Yulong Tao, Linbo Jin +10
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments…
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Qiming Shi, Yibo Dou, Jiawen Zhu +5
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summ…
Language Scent: Exploring Cross-Language Information Navigation
Jiawen Stefanie Zhu, Katharina Reinecke, Tanushree Mitra
While multilingual users often switch between languages when seeking information, this process remains undersupported by current systems where information is typically siloed by la…
Understanding Remote Communication between Grandparents and Grandchildren in Distributed Immigrant Families
Jiawen Stefanie Zhu, Jian Zhao
Grandparent-grandchild bonds are crucial for both parties. Many immigrant families are geographically dispersed, and the grandparents and grandchildren need to rely on remote commu…
Facilitating Mixed-Methods Analysis with Computational Notebooks
Jiawen Stefanie Zhu, Zibo Zhang, Jian Zhao
Data exploration is an important aspect of the workflow of mixed-methods researchers, who conduct both qualitative and quantitative analysis. However, there currently exists few to…