most citedSeeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2026

Table as a Modality for Large Language Models

Liyao Li, Chao Ye, Wentao Ye +9

To migrate the remarkable successes of Large Language Models (LLMs), the community has made numerous efforts to generalize them to the table reasoning tasks for the widely deployed…

cs.LG2025

An Invariant Latent Space Perspective on Language Model Inversion

Wentao Ye, Jiaqi Hu, Haobo Wang +7

Language model inversion (LMI), i.e., recovering hidden prompts from outputs, emerges as a concrete threat to user privacy and system security. We recast LMI as reusing the LLM's o…

cs.CL2025

CYCLE-INSTRUCT: Fully Seed-Free Instruction Tuning via Dual Self-Training and Cycle Consistency

Zhanming Shen, Hao Chen, Yulei Tang +6

Instruction tuning is vital for aligning large language models (LLMs) with human intent, but current methods typically rely on costly human-annotated seed data or powerful external…

cs.CV2025

SPA++: Generalized Graph Spectral Alignment for Versatile Domain Adaptation

Zhiqing Xiao, Haobo Wang, Xu Lu +3

Domain Adaptation (DA) aims to transfer knowledge from a labeled source domain to an unlabeled or sparsely labeled target domain under domain shifts. Most prior works focus on capt…

cs.CV20251 cited

Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

Lin Long, Yichen He, Wentao Ye +5

We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update…