7 papers
Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
Xiaofeng Shi, Xiaosong Qiu, Wenxin Ma +6
Enterprise question answering requires models to acquire proprietary knowledge without discarding general capabilities. We present Wnuan, a three-stage pipeline that constructs tas…
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
Qian Kou, Xiaofeng Shi, Yulin Li +4
Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. However, they remain brittle on mechanical eng…
ChartWalker: Benchmarking the Cross-Chart RAG Task with Hierarchical Knowledge Graphs
Ning Tang, Chenghan Xie, Hanyang Yuan +6
Cross-Chart Retrieval-Augmented Generation (RAG) is critical for complex multi-modal analytical tasks in scientific, business, and political domains. However, existing benchmarks e…
RAFT: Data Refinement and Adaptive Distillation for Domain Fine-Tuning with Alleviated Forgetting
Yuduo Li, Xiaofeng Shi, Qian Kou +2
Domain-specific supervised fine-tuning (SFT) often improves in-domain performance at the cost of degrading a model's general capabilities. We view this degradation through two prac…
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
Xiaofeng Shi, Qian Kou, Yuduo Li +1
With the rapid advancement of Large Language Models (LLMs), the Chain-of-Thought (CoT) component has become significant for complex reasoning tasks. However, in conventional Superv…
SPAR: Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic Search
Xiaofeng Shi, Yuduo Li, Qian Kou +3
Recent advances in large language models (LLMs) have opened new opportunities for academic literature retrieval. However, existing systems often rely on rigid pipelines and exhibit…