Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasks
Vansh Kapoor, Aman Gupta, Hao Chen +3
Multi-step reasoning tasks like mathematical problem solving are vulnerable to cascading failures, where a single incorrect step leads to complete solution breakdown. Current LLM r…
cs.AI2026
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
Aman Gupta, Denny O'Shea, Fazl Barez
Large language models (LLMs) are increasingly being used for tasks where outputs shape human decisions, so it is critical to verify that their responses consistently reflect desire…