4 papers
EASE-TTT: Evidence-Aligned Selective Test-Time Training for Long-Context Question Answering
Xiaopeng Yuan, Zebin Wang, Suwen Wang +3
Long-context question answering (QA) remains challenging for smaller language models even when answer-bearing evidence is already present in the input. Existing within-context retr…
Beyond Surface Statistics: Robust Conformal Prediction for LLMs via Internal Representations
Yanli Wang, Peng Kuang, Xiaoyu Han +2
Large language models are increasingly deployed in settings where reliability matters, yet output-level uncertainty signals such as token probabilities, entropy, and self-consisten…
IMPROVE: Iterative Model Pipeline Refinement and Optimization Leveraging LLM Experts
Eric Xue, Ke Chen, Zeyi Huang +2
Large language model (LLM) agents have emerged as a promising solution to automate the workflow of machine learning, but most existing methods share a common limitation: they attem…
Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding
Yuyang Ji, Haohan Wang
Multimodal Large Language Models (MLLMs) have shown remarkable versatility but face challenges in demonstrating true visual understanding, particularly in chart reasoning tasks. Ex…