activity
20242026
most citedPet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services

1 citations · 1 across the 5 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models

Huiyuan Zheng, Houtao Zhang, Boyang Wang +2

Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-neutral scenarios largely und…

cs.CL20251 cited

Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services

Hongcheng Guo, Zheyong Xie, Shaosheng Cao +6

As interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Exist…

cs.CL2025

SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services

Hongcheng Guo, Zheyong Xie, Shaosheng Cao +5

With the increasing integration of visual and textual content in Social Networking Services (SNS), evaluating the multimodal capabilities of Large Language Models (LLMs) is crucial…

cs.CL2025

H2HTalk: Evaluating Large Language Models as Emotional Companion

Boyang Wang, Yalun Wu, Hongcheng Guo +1

As digital emotional support needs grow, Large Language Model companions offer promising authentic, always-available empathy, though rigorous evaluation lags behind model advanceme…

cs.CL2025

Redefining Machine Translation on Social Network Services with Large Language Models

Hongcheng Guo, Fei Zhao, Shaosheng Cao +8

The globalization of social interactions has heightened the need for machine translation (MT) on Social Network Services (SNS), yet traditional models struggle with culturally nuan…

cs.CL2025

Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models

Hongcheng Guo, Juntao Yao, Boyang Wang +5

Mixture-of-Experts (MoE) architectures have emerged as a promising paradigm for scaling large language models (LLMs) with sparse activation of task-specific experts. Despite their…