4 papers
Improved LLM Agents for Financial Document Question Answering
Nelvin Tan, Zian Seng, Liang Zhang +3
Large language models (LLMs) have shown impressive capabilities on numerous natural language processing tasks. However, LLMs still struggle with numerical question answering for fi…
AutoMetrics: Approximate Human Judgements with Automatically Generated Evaluators
Michael J. Ryan, Yanzhe Zhang, Amol Salunkhe +3
Evaluating user-facing AI applications remains a central challenge, especially in open-ended domains such as travel planning, clinical note generation, or dialogue. The gold standa…
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
Nelvin Tan, James Asikin Cheung, Yu-Ching Shih +2
Large language models (LLMs) are becoming useful in many domains due to their impressive abilities that arise from large training datasets and large model sizes. More recently, the…
Flexible and Efficient Drift Detection without Labels
Nelvin Tan, Yu-Ching Shih, Dong Yang +1
Machine learning models are being increasingly used to automate decisions in almost every domain, and ensuring the performance of these models is crucial for ensuring high quality…