activity
20122024
most citedParametric Return Density Estimation for Reinforcement Learning

42 citations · 73 across the 12 of their papers we have counts for

collaborators

12 papers

cs.CL20241 cited

AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses

Xiaotian Lu, Jiyi Li, Koh Takeuchi +1

Question answering (QA) tasks have been extensively studied in the field of natural language processing (NLP). Answers to open-ended questions are highly diverse and difficult to q…

cs.LG202411 cited

A Generalized Model for Multidimensional Intransitivity

Jiuding Duan, Jiyi Li, Yukino Baba +1

Intransitivity is a critical issue in pairwise preference modeling. It refers to the intransitive pairwise preferences between a group of players or objects that potentially form a…

cs.LG2024

Online Policy Learning from Offline Preferences

Guoxi Zhang, Han Bao, Hisashi Kashima

In preference-based reinforcement learning (PbRL), a reward function is learned from a type of human feedback called preference. To expedite preference collection, recent works hav…

cs.LG2023

Estimating Treatment Effects Under Heterogeneous Interference

Xiaofeng Lin, Guoxi Zhang, Xiaotian Lu +3

Treatment effect estimation can assist in effective decision-making in e-commerce, medicine, and education. One popular application of this estimation lies in the prediction of the…

cs.LG2023

Label Selection Approach to Learning from Crowds

Kosuke Yoshimura, Hisashi Kashima

Supervised learning, especially supervised deep learning, requires large amounts of labeled data. One approach to collect large amounts of labeled data is by using a crowdsourcing…

cs.HC20231 cited

Mitigating Voter Attribute Bias for Fair Opinion Aggregation

Ryosuke Ueda, Koh Takeuchi, Hisashi Kashima

The aggregation of multiple opinions plays a crucial role in decision-making, such as in hiring and loan review, and in labeling data for supervised learning. Although majority vot…