22 papers
ET-SAM: Efficient Point Prompt Prediction in SAM for Unified Scene Text Detection and Layout Analysis
Xike Zhang, Maoyuan Ye, Juhua Liu +1
Previous works based on Segment Anything Model (SAM) have achieved promising performance in unified scene text detection and layout analysis. However, the typical reliance on pixel…
SkillCAT: Contrastive, Assessment-Augmented and Topology-AwareSkill Self-Evolution for LLM Agents
Kunfeng Chen, Qihuang Zhong, Juhua Liu +1
Skill self-evolution methods for LLM agents aim to turn execution trajectories into reusable skill documents. However, current pipelines typically derive skill patches from a singl…
ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering
Yikai Zhu, Kunfeng Chen, Qihuang Zhong +2
Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA), which requires reasoning o…
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
Qihuang Zhong, Liang Ding, Juhua Liu +3
Self-improvement training enables the large reasoning models (LRMs) to improve themselves by self-generating reasoning trajectories as training data without external supervision. H…
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
Shuo Ni, Tong Wang, Jing Zhang +4
Vision-Language Models (VLMs) increasingly operate on ultra-high-resolution (UHR) Earth observation imagery, yet they remain vulnerable to a severe scale mismatch between large-sca…
Learn to Think: Improving Multimodal Reasoning through Vision-Aware Self-Improvement Training
Qihuang Zhong, Liang Ding, Wenjie Xuan +3
Post-training with explicit reasoning traces is common to improve the reasoning capabilities of Multimodal Large Language Models (MLLMs). However, acquiring high-quality reasoning…