5 papers
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Shawn Li, Wei Yang, Jike Zhong +11
The paper introduces JigShape, a benchmark of interlocking jigsaw puzzles designed to test visual‑geometric reasoning in vision‑language models, and shows that current zero‑shot an…
ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling
Heng Ping, Arijit Bhattacharjee, Peiyu Zhang +5
Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents into layered reasoning pipelines. However, existing MoA variants fail to sust…
RaMem: Contextual Reinstatement for Long-term Agentic Memory
Wei Yang, Bryce Kan, Shixuan Li +5
Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems have made past experie…
Conversational Time Series Foundation Models: Towards Explainable and Effective Forecasting
Defu Cao, Michael Gee, Jinbo Liu +4
The proliferation of time series foundation models has created a landscape where no single method achieves consistent superiority, framing the central challenge not as finding the…
OMGM: Orchestrate Multiple Granularities and Modalities for Efficient Multimodal Retrieval
Wei Yang, Jingjing Fu, Rui Wang +3
Vision-language retrieval-augmented generation (RAG) has become an effective approach for tackling Knowledge-Based Visual Question Answering (KB-VQA), which requires external knowl…