activity
20222025
most citedTrustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

73 citations · 79 across the 9 of their papers we have counts for

collaborators

9 papers

cs.MA2025

How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs

Andrew Estornell, Jean-Francois Ton, Muhammad Faaiz Taufiq +1

Large Language Models (LLMs) have achieved strong performance on a wide range of complex reasoning tasks, yet further gains are often possible by leveraging the complementary stren…

stat.ML2025

Uncertainty Quantification and Causal Considerations for Off-Policy Decision Making

Muhammad Faaiz Taufiq

Off-policy evaluation (OPE) is a critical challenge in robust decision-making that seeks to assess the performance of a new policy using data collected under a different policy. Ho…

cs.CL2024★ 1 cited

Understanding Chain-of-Thought in LLMs through Information Theory

Jean-Francois Ton, Muhammad Faaiz Taufiq, Yang Liu

Large Language Models (LLMs) have shown impressive performance in complex reasoning tasks through the use of Chain-of-Thought (CoT) reasoning, allowing models to break down problem…

stat.ML2024

Achievable Fairness on Your Data With Utility Guarantees

Muhammad Faaiz Taufiq, Jean-Francois Ton, Yang Liu

In machine learning fairness, training models that minimize disparity across different sensitive groups often leads to diminished accuracy, a phenomenon known as the fairness-accur…

stat.ML2023★ 1 cited

Marginal Density Ratio for Off-Policy Evaluation in Contextual Bandits

Muhammad Faaiz Taufiq, Arnaud Doucet, Rob Cornish +1

Off-Policy Evaluation (OPE) in contextual bandits is crucial for assessing new policies using existing data without costly experimentation. However, current OPE methods, such as In…

cs.AI2023★ 73 cited

Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment

Yang Liu, Yuanshun Yao, Jean-Francois Ton +6

Ensuring alignment, which refers to making models behave in accordance with human intentions [1,2], has become a critical task before deploying large language models (LLMs) in real…