works on

From the 1 of 33 linked papers with an AI index.

activity
20242026
most citedShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents

1 citations · 1 across the 18 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation

Qicheng Zhao, Qi Sun, Zheyu Yan

The paper presents Seer, a training‑free approach that detects the true end of generated sequences in diffusion multimodal large language models by monitoring MLP activation sparsi…

cs.AI2026

ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration

Qicheng Zhao, Yu Li, Qi Sun +1

The adoption of powerful diffusion models is hindered by their significant inference latency. Recent ``cache-then-forecast'' schemes alleviate this issue by accelerating DiTs using…

cs.AI2026

From "Weak" Signals to Strong Models: Preference Delta Aggregation with LoRA Merging

Qi Sun, Siyue Zhang, Yulin Chen +3

Training strong large language models (LLMs) requires high-quality supervision, which is often scarce. Recent work shows that paired preference data from weak-weaker model pairs (e…

cs.AI2026

FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory

Yingjie Gu, Wenjian Xiong, Liqiang Wang +7

For LLM agents, memory management critically impacts efficiency, quality, and security. While much research focuses on retention, selective forgetting--inspired by human cognitive…

cs.AI2026

Search, Do not Guess: Teaching Small Language Models to Be Effective Search Agents

Yizhou Liu, Qi Sun, Yulin Chen +2

Agents equipped with search tools have emerged as effective solutions for knowledge-intensive tasks. While Large Language Models (LLMs) exhibit strong reasoning capabilities, their…

cs.AI2024

Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models

Yew Ken Chia, Qi Sun, Lidong Bing +1

Large multimodal models have demonstrated impressive problem-solving abilities in vision and language tasks, and have the potential to encode extensive world knowledge. However, it…