10 papers
Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models
Huiyuan Zheng, Houtao Zhang, Boyang Wang +2
Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-neutral scenarios largely und…
PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
Yalun Wu, Haotian Liu, Zhoujun Li +1
As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably re…
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
Anjie Le, Henan Liu, Yue Wang +18
Ultrasound is a widely-used imaging modality critical to global healthcare, yet its interpretation remains challenging due to its varying image quality on operators, noises, and an…
Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services
Hongcheng Guo, Zheyong Xie, Shaosheng Cao +6
As interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Exist…
SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services
Hongcheng Guo, Zheyong Xie, Shaosheng Cao +5
With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial…
RedOne: Revealing Domain-specific LLM Post-Training in Social Networking Services
Fei Zhao, Chonggang Lu, Yue Wang +22
As a primary medium for modern information dissemination, social networking services (SNS) have experienced rapid growth, which has proposed significant challenges for platform con…