10 papers
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Haoyu Chen, Kaichen Zhou, Hang Hua +11
The paper introduces MemoBench, a benchmark that tests video generation models' ability to remember and correctly update objects that disappear and later reappear in dynamically ch…
BEAVER: An Enterprise Benchmark for Text-to-SQL
Peter Baile Chen, Devin Yang, Weiyue Li +6
Existing text-to-SQL benchmarks have largely been constructed from public databases with well-structured schemas and simplistic question-SQL pairs. While large language models (LLM…
MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts
Weiyue Li, Ruizhi Qian, Yi Li +5
Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific conclusions from structured bio…
Do Emotions in Prompts Matter? Effects of Emotional Framing on Large Language Models
Minda Zhao, Yutong Yang, Chufei Peng +5
Emotional tone is pervasive in human communication, yet its influence on large language model (LLM) behaviour remains unclear. Here, we examine how first-person emotional framing i…
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
Dachuan Zhao, Weiyue Li, Zhenda Shen +4
Vision-Language Models (VLMs) have become indispensable for multimodal reasoning, yet their representations often encode and amplify demographic biases, resulting in biased associa…
TrustTrade: Human-Inspired Selective Consensus Reduces Decision Uncertainty in LLM Trading Agents
Minghan Li, Rachel Gonsalves, Weiyue Li +2
Large language models (LLMs) are increasingly deployed as autonomous agents in financial trading. However, they often exhibit a hazardous behavioral bias that we term uniform trust…