From the 1 of 29 linked papers with an AI index.
10 citations · 13 across the 10 of their papers we have counts for
29 papers
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
Vu Duc Anh, Nhat M. Hoang, Do Xuan Long +3
Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language models (LLMs). In this work, we propose…
TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding
Quynh Vo, Cong-Duy Nguyen, Ponhvoan Srey +2
The paper introduces TIGER, a framework that speeds up multimodal generation by dynamically selecting only the visual tokens relevant to the current textual context and training th…
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction
Mingzhe Du, Luu Anh Tuan, Tianyi Wu +4
Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar that reaches a vulnerable path,…
A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
Wenyi Xiao, Zechuan Wang, Leilei Gan +9
With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) has…
Graph-R1: Towards Agentic GraphRAG Framework via End-to-end Reinforcement Learning
Haoran Luo, Haihong E, Guanting Chen +8
Retrieval-Augmented Generation (RAG) mitigates hallucination in LLMs by incorporating external knowledge, but relies on chunk-based retrieval that lacks structural semantics. Graph…
Towards Interpretable Federated Learning
Anran Li, Rui Liu, Ming Hu +4
Federated learning (FL) enables multiple data owners to build machine learning models collaboratively without exposing their private local data. In order for FL to achieve widespre…