6 papers
GraphVerse: A Comprehensive Visual Graph Reasoning Benchmark for Multimodal Large Language Models
Yuanfu Sun, Yuanhang Ren, Kang Li +5
Recent Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse vision-language tasks, creating an urgent need for more challenging benchmarks. Yet…
Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval
Jiaxi Li, Ke Deng, Yun Wang +5
Language agents increasingly rely on reusable skills to improve multi-step web automation across related tasks. A growing line of work studies online skill learning, where agents c…
MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information
Jiaxi Li, Yucheng Shi, Xiao Huang +2
Tree search has become as a representative framework for test-time reasoning with large language models (LLMs), exemplified by methods such as Tree-of-Thought and Monte Carlo Tree…
Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
Wenqi Liu, Xuemeng Song, Jiaxi Li +4
Direct Preference Optimization (DPO) has emerged as an effective approach for mitigating hallucination in Multimodal Large Language Models (MLLMs). Although existing methods have a…
Fact or Guesswork? Evaluating Large Language Models' Medical Knowledge with Structured One-Hop Judgments
Jiaxi Li, Yiwei Wang, Kai Zhang +5
Large language models (LLMs) have been widely adopted in various downstream task domains. However, their abilities to directly recall and apply factual medical knowledge remains un…
Automating Expert-Level Medical Reasoning Evaluation of Large Language Models
Shuang Zhou, Wenya Xie, Jiaxi Li +16
As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluat…