2 citations · 2 across the 3 of their papers we have counts for
12 papers
CultureScore: Evaluating Cultural Faithfulness in Video Generation Models
Anku Rani, Wei Dai, Shravan Nayak +3
As video generation models like Veo 3.1 and LTX-2 advance, their ability to accurately represent diverse global cultures remains a critical yet understudied frontier. Current metri…
EQPO: Equitable Group Relative Policy Optimization for Clinical Reasoning
Shiqi Dai, Wei Dai, Jiaee Cheong +1
Medical AI systems demonstrated impressive diagnostic performance, yet they routinely show uneven accuracy across demographic groups, disadvantaging underrepresented populations. A…
OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization
Keane Ong, Sabri Boughorbel, Luwei Xiao +9
Socially intelligent AI systems must reason across diverse human behavioral tasks and generalize to new social contexts. However, behavioral data is inherently heterogeneous, compr…
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
Hengzhi Li, Justin Zhang, Brendon Jiang +9
Puzzlehunts are a genre of complex, multi-step puzzles lacking well-defined problem definitions. In contrast to conventional reasoning benchmarks consisting of tasks with clear ins…
SmellNet: A Large-scale Dataset for Real-world Smell Recognition
Dewei Feng, Wei Dai, Carol Li +3
The ability of AI to sense and identify various substances based on their smell alone can have profound impacts on allergen detection (e.g. smelling gluten or peanuts in a cake), m…
Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding
Keane Ong, Wei Dai, Carol Li +8
Using intelligent systems to perceive psychological and social behaviors, that is, the underlying affective, cognitive, and pathological states that are manifested through observab…