activity
20232026
most citedQuantum Bayesian Optimization

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing cs.LGShow all

13 papers · 1 filter

cs.LG2026

BarrierSteer: LLM Safety via Learning Barrier Steering

Thanh Q. Tran, Arun Verma, Kiwan Wong +3

Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe content generation remains a significant…

cs.LG2025

Incentivizing Time-Aware Fairness in Data Sharing

Jiangwei Chen, Kieu Thao Nguyen Pham, Rachael Hwee Ling Sim +4

In collaborative data sharing and machine learning, multiple parties aggregate their data resources to train a machine learning model with better model performance. However, as the…

cs.LG2025

Uncovering Scaling Laws for Large Language Models via Inverse Problems

Arun Verma, Zhaoxuan Wu, Zijian Zhou +15

Large Language Models (LLMs) are large-scale pretrained models that have achieved remarkable success across diverse domains. These successes have been driven by unprecedented compl…

cs.LG2025

COBRA: Contextual Bandit Algorithm for Ensuring Truthful Strategic Agents

Arun Verma, Indrajit Saha, Makoto Yokoo +1

This paper considers a contextual bandit problem involving multiple agents, where a learner sequentially observes the contexts and the agent's reported arms, and then selects the a…

cs.LG2025

ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment

Xiaoqiang Lin, Arun Verma, Zhongxiang Dai +3

The recent success in using human preferences to align large language models (LLMs) has significantly improved their performance in various downstream tasks, such as question answe…

cs.LG2025

Active Human Feedback Collection via Neural Contextual Dueling Bandits

Arun Verma, Xiaoqiang Lin, Zhongxiang Dai +2

Collecting human preference feedback is often expensive, leading recent works to develop principled algorithms to select them more efficiently. However, these works assume that the…