activity
20242026
most citedReasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CV2026

When Rules Fall Short: Agent-Driven Discovery of Emerging Content Issues in Short Video Platforms

Chenghui Yu, Hongwei Wang, Junwen Chen +5

Trends on short-video platforms evolve at a rapid pace, with new content issues emerging every day that fall outside the coverage of existing annotation policies. However, traditio…

cs.CV20251 cited

Reasoning-Enhanced Domain-Adaptive Pretraining of Multimodal Large Language Models for Short Video Content Governance

Zixuan Wang, Yu Sun, Hongwei Wang +6

Short video platforms are evolving rapidly, making the identification of inappropriate content increasingly critical. Existing approaches typically train separate and small classif…

cs.MM2025

Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion

Yu Sun, Yin Li, Ruixiao Sun +7

Transformer-based multimodal models are widely used in industrial-scale recommendation, search, and advertising systems for content understanding and relevance ranking. Enhancing l…

cs.CL2024

IPS: In-Prompt Process Supervision for Short Video Content Moderation

Mingchao Liu, Yu Sun, Ruixiao Sun +5

Multimodal large language models (MLLMs) are effective at capturing the semantics of short video content; however, they often fail to attend to the policy-specific details required…

cs.IR2024

USM: Unbiased Survey Modeling for Limiting Negative User Experiences in Recommendation Systems

Chenghui Yu, Peiyi Li, Haoze Wu +3

Reducing negative user experiences is essential for the success of recommendation platforms. Exposing users to inappropriate content could not only adversely affect users' psycholo…

cs.CV2024

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework

Xin Dong, Sen Jia, Ming Rui Wang +4

Recently, with the emergence of recent Multimodal Large Language Model (MLLM) technology, it has become possible to exploit its video understanding capability on different classifi…