4 papers
Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification
Yutong Wu, Xiaofan Bai, Shixin Li +10
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment. Because deployed LLMs are commonly exp…
Structure-BiEval: A Self-Supervised, Dual-Track Framework for Decoupling Structure and Content in LLM Evaluation for Web Information Systems
Boxiang Zhao, Qince Li, Zhonghao Wang +4
As Large Language Models (LLMs) evolve into the core of Web-based autonomous agents and complex Web Information Systems, their ability to faithfully translate natural language into…
Bridging the Arithmetic Gap: The Cognitive Complexity Benchmark and Financial-PoT for Robust Financial Reasoning
Boxiang Zhao, Qince Li, Zhonghao Wang +3
While Large Language Models excel at semantic tasks, they face a critical bottleneck in financial quantitative reasoning, frequently suffering from "Arithmetic Hallucinations" and…
P2S: Probabilistic Process Supervision for General-Domain Reasoning Question Answering
Wenlin Zhong, Chengyuan Liu, Yiquan Wu +5
While reinforcement learning with verifiable rewards (RLVR) has advanced LLM reasoning in structured domains like mathematics and programming, its application to general-domain rea…