5 papers
The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence
Yuxuan Ye, Jun Han, Ao Hu +7
End-to-end LLM trading agents have moved quickly from research curiosity to a small ecosystem of named systems, including FinCon, FinMem, TradingAgents, FinAgent, QuantAgent, and F…
Coarse-to-Fine Open-Set Graph Node Classification with Large Language Models
Xueqi Ma, Xingjun Ma, Sarah Monazam Erfani +2
Developing open-set classification methods capable of classifying in-distribution (ID) data while detecting out-of-distribution (OOD) samples is essential for deploying graph neura…
TensorSLM: Energy-efficient Embedding Compression of Sub-billion Parameter Language Models on Low-end Devices
Mingxue Xu, Yao Lei Xu, Danilo P. Mandic
Small Language Models (SLMs, or on-device LMs) have significantly fewer parameters than Large Language Models (LLMs). They are typically deployed on low-end devices, like mobile ph…
How can representation dimension dominate structurally pruned LLMs?
Mingxue Xu, Lisa Alazraki, Danilo P. Mandic
Pruning assumes a subnetwork exists in the original deep neural network, which can achieve comparative model performance with less computation than the original. However, it is unc…
Targeted Angular Reversal of Weights (TARS) for Knowledge Removal in Large Language Models
Harry J. Davies, Giorgos Iacovides, Danilo P. Mandic
The sheer scale of data required to train modern large language models (LLMs) poses significant risks, as models are likely to gain knowledge of sensitive topics such as bio-securi…