From the 1 of 22 linked papers with an AI index.
22 papers
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
Yuejie Li, Ke Yang, Yueying Hua +4
The paper introduces SQuTR, a benchmark dataset and evaluation protocol for testing how well spoken query‑to‑text retrieval systems perform under various levels of real‑world acous…
TRACER: Early Failure Detection for Task-Oriented Dialogue
Erfan Nourbakhsh, Rocky Slavin, Ke Yang +1
Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has already gone wrong. We present…
Answer Presence Drives RAG Rewriting Gains
Yuejie Li, Yueying Hua, Ke Yang +7
Retrieval-augmented QA pipelines often route retrieved passages through an LLM \emph{rewriter} before a smaller reader, lifting F1 by tens of points on multi-hop benchmarks; this g…
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization
Saket Reddy, Ke Yang, ChengXiang Zhai
Mitigating social bias in Large Language Models (LLMs) presents a distinct alignment challenge: unlike verifiable tasks, bias lacks a single ground truth, creating a high-variance,…
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
Rui Yang, Qianhui Wu, Zhaoyang Wang +8
Open-source native GUI agents still lag behind closed-source systems on long-horizon navigation tasks. This gap stems from two limitations: a shortage of high-quality, action-align…
Commanding Humanoid by Free-form Language: A Large Language Action Model with Unified Motion Vocabulary
Zhirui Liu, Kaiyang Ji, Ke Yang +4
Enabling humanoid robots to follow free-form natural language commands is a critical step toward seamless human-robot interaction and general-purpose embodied AI. However, existing…