1 paper
Shuo Wang, Xiangyu Wang, Quanxin Wang +9
Current evaluation practices in relational learning rely heavily on flat leaderboards that average performance across heterogeneous datasets, implicitly assuming a uniform underlyi…