From the 1 of 4 linked papers with an AI index.
4 papers
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Qiming Shi, Yulong Tao, Linbo Jin +10
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments…
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Qiming Shi, Yibo Dou, Jiawen Zhu +5
The paper introduces SKILL-KD, a contrastive skill distillation framework that creates explicit textual skill patches from teacher‑student failures to iteratively improve weaker LL…
Language Scent: Exploring Cross-Language Information Navigation
Jiawen Stefanie Zhu, Katharina Reinecke, Tanushree Mitra
While multilingual users often switch between languages when seeking information, this process remains undersupported by current systems where information is typically siloed by la…
Understanding Remote Communication between Grandparents and Grandchildren in Distributed Immigrant Families
Jiawen Stefanie Zhu, Jian Zhao
Grandparent-grandchild bonds are crucial for both parties. Many immigrant families are geographically dispersed, and the grandparents and grandchildren need to rely on remote commu…