10 papers · 1 filter
MR-Bench: A Unified Benchmark for Evidence-Grounded Multimodal Metaphor Understanding
Hong Jiang, Junnan Zhu, Jingwang Huang +9
Metaphor enables the understanding of abstract concepts through cross-domain mappings while conveying affective attitudes. In multimodal scenarios, visual and textual information j…
Me-Agent: A Personalized Mobile Agent with Two-Level User Habit Learning for Enhanced Interaction
Shuoxin Wang, Chang Liu, Gowen Loo +5
Large Language Model (LLM)-based mobile agents have made significant performance advancements. However, these agents often follow explicit user instructions while overlooking perso…
MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
Kaiwen Wei, Rui Shan, Dongsheng Zou +4
Large reasoning models (LRMs) have shown significant progress in test-time scaling through chain-of-thought prompting. Current approaches like search-o1 integrate retrieval augment…
ES-Mem: Event Segmentation-Based Memory for Long-Term Dialogue Agents
Huhai Zou, Tianhao Sun, Chuanjiang He +6
Memory is critical for dialogue agents to maintain coherence and enable continuous adaptation in long-term interactions. While existing memory mechanisms offer basic storage and re…
DiffER: Diffusion Entity-Relation Modeling for Reversal Curse in Diffusion Large Language Models
Shaokai He, Kaiwen Wei, Xinyi Zeng +5
The "reversal curse" refers to the phenomenon where large language models (LLMs) exhibit predominantly unidirectional behavior when processing logically bidirectional relationships…
ReasonTabQA: A Comprehensive Benchmark for Table Question Answering from Real World Industrial Scenarios
Changzai Pan, Jie Zhang, Kaiwen Wei +15
Recent advancements in Large Language Models (LLMs) have significantly catalyzed table-based question answering (TableQA). However, existing TableQA benchmarks often overlook the i…