6 papers
Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity
Hongnan Ma, Yiwei Shi, Mengyue Yang +1
Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's prediction, but also necessary for mainta…
OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios
Xinyi Li, Zhen Fang, Yongxin Deng +12
Hallucination detection is essential for the reliable deployment of large language models (LLMs). However, existing evaluations face two core challenges: inconsistent inference con…
StakeBench: Evaluating Language Understanding Grounded in Market Commitment
Yunhua Pei, Jingyu Hu, Yiwei Shi +3
Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers have committed to in the market.…
CreativeGame:Toward Mechanic-Aware Creative Game Generation
Hongnan Ma, Han Wang, Shenglin Wang +6
Large language models can generate plausible game code, but turning this capability into \emph{iterative creative improvement} remains difficult. In practice, single-shot generatio…
Dynamic Correction of Erroneous State Estimates via Diffusion Bayesian Exploration
Yiwei Shi, Hongnan Ma, Mengyue Yang +2
In emergency response and other high-stakes societal applications, early-stage state estimates critically shape downstream outcomes. Yet, these initial state estimates-often based…
TriShGAN: Enhancing Sparsity and Robustness in Multivariate Time Series Counterfactuals Explanation
Hongnan Ma, Yiwei Shi, Guanxiong Sun +2
In decision-making processes, stakeholders often rely on counterfactual explanations, which provide suggestions about what should be changed in the queried instance to alter the ou…