From the 2 of 16 linked papers with an AI index.
16 papers
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Qiming Shi, Yulong Tao, Linbo Jin +10
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments…
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Qiming Shi, Yibo Dou, Jiawen Zhu +5
The paper introduces SKILL-KD, a contrastive skill distillation framework that creates explicit textual skill patches from teacher‑student failures to iteratively improve weaker LL…
Dependency-Guided Code Generation: Structured Matrix Decomposition and Consistency-Guided Refinement
Mingqiao Mo, Yangchen Zeng, Zikai Xiao +7
The increasing complexity of modern software systems has made automated code generation a fundamental task in software engineering. However, existing approaches often fail to adequ…
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
Eric Hanchen Jiang, Zhi Zhang, Yuchen Wu +11
The paper introduces MemCon, a framework that treats memory operations of large language model agents as a controllable Markov Decision Process, learning adaptive policies for when…
SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering
Qiming Shi, Zhaolu Kang, Yunfan Zhou +2
Large language models are increasingly deployed as tool-augmented agents to acquire information beyond parametric knowledge. While recent work has improved long-horizon tool-use re…
When Seeing Is Not Believing -- A Benchmark for Search-Grounded Video Misinformation Detection
Tao Yu, Yujia Yang, Shenghua Chai +17
Video misinformation increasingly operates at the semantic and evidential level: authentic footage may be selectively edited, temporally reordered, spliced across sources, or augme…