73 citations · 79 across the 9 of their papers we have counts for
9 papers
How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs
Andrew Estornell, Jean-Francois Ton, Muhammad Faaiz Taufiq +1
Large Language Models (LLMs) have achieved strong performance on a wide range of complex reasoning tasks, yet further gains are often possible by leveraging the complementary stren…
Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making
Muhammad Faaiz Taufiq
Off-policy evaluation (OPE) is a critical challenge in robust decision-making that seeks to assess the performance of a new policy using data collected under a different policy. Ho…
Understanding Chain-of-Thought in LLMs through Information Theory
Jean-Francois Ton, Muhammad Faaiz Taufiq, Yang Liu
Large Language Models (LLMs) have shown impressive performance in complex reasoning tasks through the use of Chain-of-Thought (CoT) reasoning, allowing models to break down problem…
Achievable Fairness on Your Data With Utility Guarantees
Muhammad Faaiz Taufiq, Jean-Francois Ton, Yang Liu
In machine learning fairness, training models that minimize disparity across different sensitive groups often leads to diminished accuracy, a phenomenon known as the fairness-accur…
Marginal Density Ratio for Off-Policy Evaluation in Contextual Bandits
Muhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish +1
Off-Policy Evaluation (OPE) in contextual bandits is crucial for assessing new policies using existing data without costly experimentation. However, current OPE methods, such as In…
Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton +6
Ensuring alignment, which refers to making models behave in accordance with human intentions [1,2], has become a critical task before deploying large language models (LLMs) in real…