#large language models
510 resultsAISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin, Shenzhe Zhu, Shu Yang +23
The paper presents AISPA, a user‑centric framework for auditing the system prompts that guide large language model behavior in commercial AI products, and reports findings from ana…
Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
Phuc Pham, Truong-Son Hy
The paper proposes a Function–Evidence–Validation (FEV) framework to evaluate bioinformatics workflows generated by large language model agents, emphasizing workflow correctness an…
LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference
Sangjin Kim, Yuseon Choi, Jungjun Oh +2
LightRot introduces a lightweight rotation scheme and a dedicated hardware accelerator that enable energy‑efficient, low‑bit inference for large language models such as LLaMA2‑13B…
CHARGE: Leveraging CWE Hierarchies for Hardware Security SystemVerilog Assertion Generation
Xiao Tan, Cynthia Sturton
CHARGE is an automated framework that uses CWE hierarchies and large language models to generate SystemVerilog assertions for unverified RTL modules, enabling security property inf…
Revisiting Predictive Process Monitoring in the Age of Foundation Models: A Comparative Study of Sequence, Tabular, and LLM Approaches
Lennart Fertig, Lukas Kirchdorfer, Tobias Sesterhenn
The paper benchmarks predictive process monitoring using three approaches—sequence models, tabular foundation models, and large language models—across several datasets and tasks, f…
STEREODISCO: Discovering Stereotypicality in LLMs
Farane Jalali Farahani, Corina Dima, Mojtaba Nayyeri +2
The paper introduces STEREODISCO, a framework that discovers stereotypical semantic axes in large language models by probing their activation spaces using WordNet antonym pairs, an…
DeepResearch Agent System
Yong Huang, Yulu Huang, for the team Collaboration
The DeepResearch Agent System is a large language model designed for deep information retrieval and multi-step autonomous research, using a sparse activation architecture that acti…
Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling
Yuan Tian, Yi Mei, Mengjie Zhang
The paper proposes using heuristic rules evolved by genetic programming to guide large language models in making dynamic multi-mode project scheduling decisions, improving performa…
ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents
Xingjian Wu, Xuhang Zhu, Xingchen Liu +6
The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Qiming Shi, Yibo Dou, Jiawen Zhu +5
The paper introduces SKILL-KD, a contrastive skill distillation framework that creates explicit textual skill patches from teacher‑student failures to iteratively improve weaker LL…
LLM-Guided Evolutionary Search for Constraint Model Reformulation to Improve Solver Efficiency
Kostis Michailidis, Dimos Tsouros, Nguyen Dang +1
The paper explores using large language models within an evolutionary search framework to automatically reformulate constraint models for faster solving, introducing a diversity‑pr…
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Bertil Braun, Martin Forell
The paper presents a scalable framework for automatically evaluating large language model outputs using pairwise comparisons and an Elo rating system, achieving rankings that align…
Beyond Sentiment: Structured Information Extraction from Financial News
Daohan Zhu, Sitong Ge, Ruofei Wang +4
The paper proposes a framework that uses a large language model to extract multiple semantic dimensions (event type, impact scope, temporal horizon, confidence, etc.) from financia…
AI systems and the reproduction of (standard) language ideologies in World Englishes
Kingsley Ugwuanyi
The paper investigates how large language models and related AI systems reproduce standard language ideologies that privilege Inner Circle English, marginalizing non‑dominant varie…
Evaluating and Pricing Advertisements in AI-Generated Responses
John L. Turner-Smith, Zimeng Huang, Yuhan Fu +2
The paper introduces a psychologically grounded agent simulation to create supervision for predicting click‑through intent of ads embedded in LLM‑generated responses, builds a ligh…
Hierarchical Latent Reasoning for LLM-based Recommendation
Peiyu Hu, Siying Gu, Weihai Lu +8
The paper introduces HiLaR, a framework that uses hierarchical latent reasoning and layer-aware reinforcement optimization to improve recommendation performance of large language m…
LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
Shuang Liang, Haoyang Zhou, Yifan Gong +2
The paper introduces LEEPS, a latent-guided explore‑exploit prompt sampler that selects prompts before rollout to reduce wasted generation budget and improve reinforcement learning…
CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance
Anubhav Lakra, Yue Feng
The paper introduces CACHE-UK, a stability-aware memory editing framework for 4-bit quantized large language models used in UK finance, which reduces knowledge degradation during s…
AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas
Lennart Meincke, Karan Girotra, Gideon Nave +3
The paper evaluates how large language models like GPT‑4 can generate new product ideas for college students, finding that AI‑generated ideas have higher predicted purchase intent…
Towards joint scaling laws with optimal batch size schedules
Jiaxiang Li, Zhiqi Bu, Shiyun Xu
The paper derives a theoretical relationship between learning rate and batch size schedules using convex optimization, and proposes a closed‑form optimal batch size schedule that i…
From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents
Xu Xia, Jinghua Piao, Min Yang +3
The paper introduces Outcome-Verified Comparative Self-Distillation (OVCSD), a method that lets large language model agents internalize skills by supervising them with teachers who…
Group-Reflective Self-Distillation for Agentic Reinforcement Learning
Binbin Zheng, Zijun Xie, Guanqun Zhao +4
The paper introduces Group-Reflective Self-Distillation (GRSD), a method that uses a policy's own verified rollouts to generate privileged guidance for better credit assignment in…
Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations
Pere Martra, Eugenio MartÃnez Cámara, Alfonso Ureña López
The paper introduces Fairness Pruning, a method that identifies and zeroes a small set of neurons in GLU-MLP layers of large language models to locate and modulate demographic bias…
Can Large Language Models Execute Parent Orders?
Zane Shen, Xinli Xu, Guangyi Zhang +7
The paper investigates using large language models for executing large parent orders in algorithmic trading, introducing a hierarchical framework called PACE that plans and execute…