5 papers
Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
Yuxuan Ye, Raul Santos-Rodriguez, Edwin Simpson
Grounded claim factuality checking is important for large language model (LLM) applications such as retrieval-augmented generation, as it helps users assess the correctness of gene…
Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics
Yuxuan Ye, Raul Santos-Rodriguez, Edwin Simpson
Reinforcement learning with evaluation metrics as rewards is widely used to enhance specific capabilities of language models. However, for tasks such as factually consistent summar…
The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence
Yuxuan Ye, Jun Han, Ao Hu +7
End-to-end LLM trading agents have moved quickly from research curiosity to a small ecosystem of named systems, including FinCon, FinMem, TradingAgents, FinAgent, QuantAgent, and F…
Quantifying Compositionality of Classic and State-of-the-Art Embeddings
Zhijin Guo, Chenhao Xue, Zhaozhen Xu +4
For language models to generalize correctly to novel expressions, it is critical that they exploit access compositional meanings when this is justified. Even if we don't know what…
Using Similarity to Evaluate Factual Consistency in Summaries
Yuxuan Ye, Edwin Simpson, Raul Santos Rodriguez
Cutting-edge abstractive summarisers generate fluent summaries, but the factuality of the generated text is not guaranteed. Early summary factuality evaluation metrics are usually…