8 papers
Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs
Polydoros Giannouris, Mohsinul Kabir, Sophia Ananiadou
LLM deception is often evaluated through direct markers such as fabricated claims, explicit lies, or strategic concealment. However, many real-world misleading communications do no…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
Moira: Language-driven Hierarchical Reinforcement Learning for Pair Trading
Polydoros Giannouris, Yuechen Jiang, Lingfei Qian +5
Many sequential decision-making problems exhibit hierarchical structure, where high-level semantic choices constrain downstream actions and feedback is delayed and ambiguous. Learn…
NOMAD: A Multi-Agent LLM System for UML Class Diagram Generation from Natural Language Requirements
Polydoros Giannouris, Sophia Ananiadou
Large Language Models (LLMs) are increasingly utilised in software engineering, yet their ability to generate structured artefacts such as UML diagrams remains underexplored. In th…
Same Claim, Different Judgment: Benchmarking Scenario-Induced Bias in Multilingual Financial Misinformation Detection
Zhiwei Liu, Yupen Cao, Yuechen Jiang +22
Large language models (LLMs) have been widely applied across various domains of finance. Since their training data are largely derived from human-authored corpora, LLMs may inherit…
Semantic Label Drift in Cross-Cultural Translation
Mohsinul Kabir, Tasnim Ahmed, Md Mezbaur Rahman +2
Machine Translation (MT) is widely employed to address resource scarcity in low-resource languages by generating synthetic data from high-resource counterparts. While sentiment pre…