papers

Publications (11)

cs.CL2026

Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks

Wenbo Pan, Jie Xu, Qiguang Chen +5

Large Language Models (LLMs) should refuse to answer questions beyond their knowledge. This capability, which we term knowledge-aware refusal, is crucial for factual reliability, w…

cs.CR2025

Exploring and Exploiting the Resource Isolation Attack Surface of WebAssembly Containers

Zhaofeng Yu, Dongyang Zhan, Lin Ye +3

Recently, the WebAssembly (or Wasm) technology has been rapidly evolving, with many runtimes actively under development, providing cross-platform secure sandboxes for Wasm modules…

cs.LG2025

Breaking the Context Bottleneck on Long Time Series Forecasting

Chao Ma, Yikai Hou, Xiang Li +4

Long-term time-series forecasting is essential for planning and decision-making in economics, energy, and transportation, where long foresight is required. To obtain such long fore…

stat.ML2025

Do Contemporary Causal Inference Models Capture Real-World Heterogeneity? Findings from a Large-Scale Benchmark

Haining Yu, Yizhou Sun

We present unexpected findings from a large-scale benchmark study evaluating Conditional Average Treatment Effect (CATE) estimation algorithms, i.e., CATE models. By running 16 mod…

cond-mat.stat-mech2020

The dual formalisms of nonextensive thermodynamics for open systems with maximum entropy principle

Yahui Zheng, Haining Yu, Jiulin Du

We study the nonextensive thermodynamics for open systems. On the basis of the maximum entropy principle, the dual power-law q-distribution functions are re-deduced by using the du…

cs.AI2026

QDA-SQL: Questions Enhanced Dialogue Augmentation for Multi-Turn Text-to-SQL

Yinggang Sun, Ziming Guo, Haining Yu +5

The paper introduces QDA-SQL, a data augmentation technique that uses large language models to generate and validate multi‑turn question‑answer pairs, improving fine‑tuned models'…

#text-to-sql#multi-turn dialogue#data augmentation#large language models
cs.LG2026

Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMs

Wenbo Pan, Zhichao Liu, Xianlong Wang +2

Token attribution methods provide intuitive explanations for language model outputs by identifying causally important input tokens. However, as modern LLMs increasingly rely on ext…

cond-mat.stat-mech2017

The nonextensive parameter for the rotating astrophysical systems with power-law distributions

Haining Yu, Jiulin Du

We study the nonextensive parameter for the rotating astrophysical systems with power-law distributions, including both the rotating self-gravitating system and the rotating space…

cs.LG2025

Long Input Sequence Network for Long Time Series Forecasting

Chao Ma, Yikai Hou, Xiang Li +2

Short fixed-length inputs are the main bottleneck of deep learning methods in long time-series forecasting tasks. Prolonging input length causes overfitting, rapidly deteriorating…

cs.CL2025

The Hidden Dimensions of LLM Alignment: A Multi-Dimensional Analysis of Orthogonal Safety Directions

Wenbo Pan, Zhichao Liu, Qiguang Chen +3

Large Language Models' safety-aligned behaviors, such as refusing harmful queries, can be represented by linear directions in activation space. Previous research modeled safety beh…

cs.CR2026

WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation

Zhichao Liu, Wenbo Pan, Haining Yu +3

Browser agents are increasingly deployed in long-horizon tasks, which require executing extended action chains to accomplish user goals. However, this prolonged execution process p…