21 papers
Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents
Wei Shao, Chongzhou Fang, Zuxiong Tan +4
Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend on services, state, or access discovered…
Unit-Independent Low-Rate Wrist GSR Processing for Stress Detection Using Phasic nSCR Features
Zequan Liang, Sally Hang, Geneva M. Jost +8
Galvanic skin response (GSR) is widely used for stress detection, but wrist-based GSR remains challenging because its absolute amplitude can differ substantially from laboratory-gr…
Low-Rate Wrist SpO2 Estimation under Micro-Perturbations Using Motion-Aware Beat Selection and Perfusion-Guided Calibration
Zequan Liang, Ning Miao, Wei Shao +5
Continuous oxygen saturation (SPO2) monitoring from photoplethysmography (PPG) is important for wearable health sensing, but wrist-based SPO2 estimation remains challenging due to…
SpO Predictor-Guided Stage-Wise Time-Frequency Reconstruction of Low-Quality Dual-Wavelength PPG for Oxygen Saturation Estimation
Zequan Liang, Elahe Hosseini, Ning Miao +5
Continuous oxygen saturation (SpO) estimation from wearable photoplethysmography (PPG) is important for long-term health monitoring, but low-quality red and infrared PPG segmen…
STABLEVAL: Disagreement-Aware and Stable Evaluation of AI Systems
Akash Bonagiri, Gerard Janno Anderias, Saee Patil +6
Human evaluation remains the primary standard for assessing modern AI systems, yet annotator disagreement, bias, and variability make system rankings fragile under standard majorit…
CausalFlow: Causal Attribution and Counterfactual Repair for LLM Agent Failures
Akash Bonagiri, Devang Borkar, Gerard Janno Anderias +2
Large language model (LLM) agents frequently fail on multi-step tasks involving reasoning, tool use, and environment interaction. While such failures are typically logged or retrie…