3 papers
cs.CL2026
LatCom: Cross-Agent Latent Compression for Efficient Multi-Agent Collaboration
Shinan Zhang, Tao Zhang, Qihui Zhu +7
LLM-based multi-agent systems (MAS) increasingly use latent collaboration to avoid the information loss and repeated encoding-decoding overhead of natural-language communication. H…
cs.CV2026
GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
Mengjie Zhang, Qihui Zhu, Tao Zhang +10
Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number of redundant spatio-temporal v…
cs.CL2026
Rethinking Table Pruning in TableQA: From Sequential Revisions to Gold Trajectory-Supervised Parallel Search
Yu Guo, Shenghao Ye, Shuangwu Chen +8
Table Question Answering (TableQA) benefits significantly from table pruning, which extracts compact sub-tables by eliminating redundant cells to streamline downstream reasoning. H…