NewEvery arXiv paper, its researchers & institutions — mapped.
the archive

#large language models

510 results
cs.AI2026

AISPA: User-Centric System Prompt Auditing for Large Language Model Applications

Xiangning Lin, Shenzhe Zhu, Shu Yang +23

The paper presents AISPA, a user‑centric framework for auditing the system prompts that guide large language model behavior in commercial AI products, and reports findings from ana…

#system prompts#large language models#audit framework#user protection
cs.AI2026

Evaluating Agentic Bioinformatics through Function, Evidence, and Validation

Phuc Pham, Truong-Son Hy

The paper proposes a Function–Evidence–Validation (FEV) framework to evaluate bioinformatics workflows generated by large language model agents, emphasizing workflow correctness an…

#bioinformatics#large language models#agentic systems#workflow evaluation
cs.AR2026

LightRot: A Light-Weighted Rotation Scheme and Architecture for Accurate Low-Bit Large Language Model Inference

Sangjin Kim, Yuseon Choi, Jungjun Oh +2

LightRot introduces a lightweight rotation scheme and a dedicated hardware accelerator that enable energy‑efficient, low‑bit inference for large language models such as LLaMA2‑13B…

#low-bit quantization#large language models#hardware accelerator#fast hadamard transform
cs.CR2026

CHARGE: Leveraging CWE Hierarchies for Hardware Security SystemVerilog Assertion Generation

Xiao Tan, Cynthia Sturton

CHARGE is an automated framework that uses CWE hierarchies and large language models to generate SystemVerilog assertions for unverified RTL modules, enabling security property inf…

#systemverilog assertions#cwe hierarchy#large language models#hardware security verification
cs.LG2026

Revisiting Predictive Process Monitoring in the Age of Foundation Models: A Comparative Study of Sequence, Tabular, and LLM Approaches

Lennart Fertig, Lukas Kirchdorfer, Tobias Sesterhenn

The paper benchmarks predictive process monitoring using three approaches—sequence models, tabular foundation models, and large language models—across several datasets and tasks, f…

#predictive process monitoring#sequence models#tabular foundation models#large language models
cs.AI2026

STEREODISCO: Discovering Stereotypicality in LLMs

Farane Jalali Farahani, Corina Dima, Mojtaba Nayyeri +2

The paper introduces STEREODISCO, a framework that discovers stereotypical semantic axes in large language models by probing their activation spaces using WordNet antonym pairs, an…

#stereotype detection#large language models#semantic axes#probing
cs.AI2026

DeepResearch Agent System

Yong Huang, Yulu Huang, for the team Collaboration

The DeepResearch Agent System is a large language model designed for deep information retrieval and multi-step autonomous research, using a sparse activation architecture that acti…

#large language models#information retrieval#multi-step reasoning#agent systems
cs.AI2026

Guiding Large Language Models with Genetic Programming-Evolved Heuristic Knowledge for Dynamic Multi-Mode Project Scheduling

Yuan Tian, Yi Mei, Mengjie Zhang

The paper proposes using heuristic rules evolved by genetic programming to guide large language models in making dynamic multi-mode project scheduling decisions, improving performa…

#project scheduling#genetic programming#large language models#heuristic optimization
cs.LG2026

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

Xingjian Wu, Xuhang Zhu, Xingchen Liu +6

The paper introduces ClawTrack, a benchmark that evaluates both the final outcomes and the step-by-step reasoning processes of LLM-based autonomous agents across multiple dimension…

#large language models#autonomous agents#benchmarking#process evaluation
cs.AI2026

SKILL-KD: Contrastive Skill Distillation for LLM Agents

Qiming Shi, Yibo Dou, Jiawen Zhu +5

The paper introduces SKILL-KD, a contrastive skill distillation framework that creates explicit textual skill patches from teacher‑student failures to iteratively improve weaker LL…

#large language models#skill distillation#contrastive learning#agent adaptation
cs.AI2026

LLM-Guided Evolutionary Search for Constraint Model Reformulation to Improve Solver Efficiency

Kostis Michailidis, Dimos Tsouros, Nguyen Dang +1

The paper explores using large language models within an evolutionary search framework to automatically reformulate constraint models for faster solving, introducing a diversity‑pr…

#constraint programming#large language models#evolutionary algorithms#model reformulation
cs.CL2026

(Towards) Scalable Reliable Automated Evaluation with Large Language Models

Bertil Braun, Martin Forell

The paper presents a scalable framework for automatically evaluating large language model outputs using pairwise comparisons and an Elo rating system, achieving rankings that align…

#automated evaluation#large language models#pairwise comparison#elo rating
cs.CL2026

Beyond Sentiment: Structured Information Extraction from Financial News

Daohan Zhu, Sitong Ge, Ruofei Wang +4

The paper proposes a framework that uses a large language model to extract multiple semantic dimensions (event type, impact scope, temporal horizon, confidence, etc.) from financia…

#financial news#sentiment analysis#structured information extraction#event detection
cs.CL2026

AI systems and the reproduction of (standard) language ideologies in World Englishes

Kingsley Ugwuanyi

The paper investigates how large language models and related AI systems reproduce standard language ideologies that privilege Inner Circle English, marginalizing non‑dominant varie…

#world englishes#language ideology#large language models#standardization
cs.AI2026

Evaluating and Pricing Advertisements in AI-Generated Responses

John L. Turner-Smith, Zimeng Huang, Yuhan Fu +2

The paper introduces a psychologically grounded agent simulation to create supervision for predicting click‑through intent of ads embedded in LLM‑generated responses, builds a ligh…

#advertising#large language models#click-through prediction#pricing mechanisms
cs.IR2026

Hierarchical Latent Reasoning for LLM-based Recommendation

Peiyu Hu, Siying Gu, Weihai Lu +8

The paper introduces HiLaR, a framework that uses hierarchical latent reasoning and layer-aware reinforcement optimization to improve recommendation performance of large language m…

#recommendation systems#large language models#latent reasoning#hierarchical representation
cs.CL2026

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

Shuang Liang, Haoyang Zhou, Yifan Gong +2

The paper introduces LEEPS, a latent-guided explore‑exploit prompt sampler that selects prompts before rollout to reduce wasted generation budget and improve reinforcement learning…

#prompt sampling#reinforcement learning#large language models#explore-exploit
cs.CL2026

CACHE-UK: A Stability-Aware Memory Editor for Sequentially Updated Quantized LLMs in Finance

Anubhav Lakra, Yue Feng

The paper introduces CACHE-UK, a stability-aware memory editing framework for 4-bit quantized large language models used in UK finance, which reduces knowledge degradation during s…

#large language models#quantization#memory editing#financial domain
cs.AI2026

AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas

Lennart Meincke, Karan Girotra, Gideon Nave +3

The paper evaluates how large language models like GPT‑4 can generate new product ideas for college students, finding that AI‑generated ideas have higher predicted purchase intent…

#large language models#idea generation#product design#creativity
cs.LG2026

Towards joint scaling laws with optimal batch size schedules

Jiaxiang Li, Zhiqi Bu, Shiyun Xu

The paper derives a theoretical relationship between learning rate and batch size schedules using convex optimization, and proposes a closed‑form optimal batch size schedule that i…

#dynamic batch size#learning rate schedule#scaling laws#convex optimization
cs.AI2026

From Scoring to Acting: Outcome-Verified Comparative Self-Distillation for LLM Agents

Xu Xia, Jinghua Piao, Min Yang +3

The paper introduces Outcome-Verified Comparative Self-Distillation (OVCSD), a method that lets large language model agents internalize skills by supervising them with teachers who…

#large language models#self-distillation#reinforcement learning#agent training
cs.AI2026

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

Binbin Zheng, Zijun Xie, Guanqun Zhao +4

The paper introduces Group-Reflective Self-Distillation (GRSD), a method that uses a policy's own verified rollouts to generate privileged guidance for better credit assignment in…

#agentic reinforcement learning#self-distillation#verifiable rewards#large language models
cs.CL2026

Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

Pere Martra, Eugenio Martínez Cámara, Alfonso Ureña López

The paper introduces Fairness Pruning, a method that identifies and zeroes a small set of neurons in GLU-MLP layers of large language models to locate and modulate demographic bias…

#bias mitigation#large language models#neuron interpretability#GLU architecture
cs.CE2026

Can Large Language Models Execute Parent Orders?

Zane Shen, Xinli Xu, Guangyi Zhang +7

The paper investigates using large language models for executing large parent orders in algorithmic trading, introducing a hierarchical framework called PACE that plans and execute…

#algorithmic trading#parent order execution#large language models#hierarchical planning