8 papers
IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning
Wei Zhang, Xinwu Liu, Yihang Cheng
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT) of LLMs whose effectiveness depends on rank allocation. Existing adaptive LoRA method…
Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents
Kexin Chu, Dawei Xiang, Wei Zhang
Tool-augmented LLM agents increasingly access the same tool type through multiple functionally equivalent providers, such as web-search APIs, retrievers, or LLM backends exposed be…
MemQ: Integrating Q-Learning into Self-Evolving Memory Agents over Provenance DAGs
Junwei Liao, Haoting Shi, Ruiwen Zhou +9
Episodic memory allows LLM agents to accumulate and retrieve experience, but current methods treat each memory independently, i.e., evaluating retrieval quality in isolation withou…
Probing Routing-Conditional Calibration in Attention-Residual Transformers
Wenhao Liang, Lin Yue, Wei Emma Zhang +4
Post-hoc calibration is usually evaluated as a function of logits or softmax confidence alone, even as routing-augmented architectures increasingly accompany predictions with sampl…
Explainable Knowledge Tracing via Probabilistic Embeddings and Pattern-based Reasoning
Siyu Wu, Cong Xu, Wei Zhang
Knowledge Tracing (KT) models students' knowledge states based on learning interactions to predict performance. While deep learning-based KT models have boosted predictive accuracy…
SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation
Huimin Xu, Xin Mao, Feng-Lin Li +4
Process Reward Models (PRMs) have demonstrated promising results in mathematical reasoning, but existing process annotation approaches, whether through human annotations or Monte C…