8 papers
Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
Mingchen Li, Meikang Qiu, Zifan Peng +4
Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for legitimate code review, triage, and repair…
"What Did It Actually Do?": Understanding Risk Awareness and Traceability for Computer-Use Agents
Zifan Peng, Mingchen Li
Personalized computer-use agents are rapidly moving from expert communities into mainstream use. Unlike conventional chatbots, these systems can install skills, invoke tools, acces…
Enhanced Structured Lasso Pruning with Class-wise Information
Xiang Liu, Mingchen Li, Xia Li +7
Modern applications require lightweight neural network models. Most existing neural network pruning methods focus on removing unimportant filters; however, these may result in the…
TxSum: User-Centered Ethereum Transaction Understanding with Micro-Level Semantic Grounding
Zifan Peng, Jingyi Zheng, Yule Liu +8
Understanding the economic intent of Ethereum transactions is critical for user safety, yet current tools expose only raw on-chain data or surface-level intent, leading to widespre…
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
Zifan Peng, Yule Liu, Zhen Sun +9
Large Audio Language Models (LALMs) have made significant progress. While increasingly deployed in real-world applications, LALMs face growing safety risks from jailbreak attacks t…
NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
Jingzhe Ding, Shengda Long, Changxin Pu +46
Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks fail to rigorously evaluate the long-horizon capabilities re…