6 papers
Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
Cheng Fan, Junyi Zhou, Tingzhang Luo +5
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering h…
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
Qiyanhui Lu, Han Wu, Rongjian Xu +6
Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing training-free pruning methods sele…
TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding
Chengcheng Wang, Tingzhang Luo, Wenhao Li +2
Diffusion language models (DLLMs) generate text by iteratively denoising masked positions, exposing a trajectory of predictive distributions rather than a single instantaneous beli…
From Question Answering to Task Completion: A Survey on Agent System and Harness Design
Jianyuan Guo, Zhiwei Hao, Chengcheng Wang +14
LLM-based agents mark a shift from passive question answering to active task completion: they perceive environments, invoke tools, maintain state, and act over extended horizons. A…
An Ensemble of Evolutionary Algorithms With Both Crisscross Search and Sparrow Search for Processing Inferior Individuals
Mingxuan Du, Tingzhang Luo, Ziyang Wang +1
In the field of artificial intelligence, real parameter single objective optimization is an important direction. Both the Differential Evolution (DE) and the Covariance Matrix Adap…
DIG-FACE: De-biased Learning for Generalized Facial Expression Category Discovery
Tingzhang Luo, Yichao Liu, Yuanyuan Liu +6
We introduce a novel task, Generalized Facial Expression Category Discovery (G-FACE), that discovers new, unseen facial expressions while recognizing known categories effectively.…