From the 1 of 71 linked papers with an AI index.
6 citations · 6 across the 24 of their papers we have counts for
43 papers · 1 filter
From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models
Ponhvoan Srey, Xiaobao Wu, Cong-Duy Nguyen +3
Probe-based uncertainty estimation (UE) has emerged as a prominent approach to detect hallucinations in Large Language Models (LLMs) by learning uncertainty from internal model sig…
EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
Jundong Xu, Qingchuan Li, Jiaying Wu +11
Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deploymen…
GRACE: Step-Level Benchmark for Faithful Reasoning over Context
Hoang Pham, Dong Le, Anh Tuan Luu
Many reasoning tasks require models to reason over input context, from document-grounded question answering to rule-based deduction. Chain-of-Thought (CoT) prompting produces trace…
Don't Read Everything: A Curvature-Conditioned Query for Linear Attention
Dong Le, Thong Nguyen, Cong-Duy Nguyen +1
Linear attention reduces the quadratic cost of softmax attention by maintaining a recurrent fast-weight state, but it consistently lags on in-context retrieval and long-context tas…
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
Khoi Le, Tri Cao, Phong Nguyen +5
Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train-test distributions. Therefore,…
Gradient-Boosted Decision Tree for Listwise Context Model in Multimodal Review Helpfulness Prediction
Thong Nguyen, Xiaobao Wu, Xinshuai Dong +4
Multimodal Review Helpfulness Prediction (MRHP) aims to rank product reviews based on predicted helpfulness scores and has been widely applied in e-commerce via presenting customer…