13 papers
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty +2
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation…
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
Zhuohan Xie, Daniil Orel, Rushil Thareja +22
Multi-step symbolic reasoning is essential for robust financial analysis; yet, current benchmarks largely overlook this capability. Existing datasets such as FinQA and ConvFinQA em…
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
Dhruv Sahnan, Subhabrata Dutta, Tanmoy Chakraborty +2
Professional fact-checkers rely on domain knowledge and deep contextual understanding to verify claims. Large language models (LLMs) and large reasoning models (LRMs) lack such gro…
The CLEF-2026 CheckThat! Lab: Advancing Multilingual Fact-Checking
Julia Maria StruÃ, Sebastian Schellhammer, Stefan Dietze +9
The CheckThat! lab aims to advance the development of innovative technologies combating disinformation and manipulation efforts in online communication across a multitude of langua…
Can LLMs Automate Fact-Checking Article Writing?
Dhruv Sahnan, David Corney, Irene Larraz +7
Automatic fact-checking aims to support professional fact-checkers by offering tools that can help speed up manual fact-checking. Yet, existing frameworks fail to address the key s…
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
Aishik Mandal, Tanmoy Chakraborty, Iryna Gurevych
Mental health disorders create profound personal and societal burdens, yet conventional diagnostics are resource-intensive and limit accessibility. Advances in artificial intellige…