3 papers
cs.MM2026
UNIVID: Unified Vision-Language Model for Video Moderation
Kejuan Yang, Yizhuo Zhang, Mingyuan Du +6
Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to support downstream enforcement. Tr…
cs.HC2026
GazeSummary: Exploring Gaze as an Implicit Prompt for Personalization in Text-based LLM Tasks
Jiexin Ding, Yizhuo Zhang, Xinyun Liu +4
Smart glasses are accelerating progress toward more seamless and personalized LLM-based assistance by integrating multimodal inputs. Yet, these inputs rely on obtrusive explicit pr…
cs.LG2025
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
Yizhuo Zhang, Heng Wang, Shangbin Feng +3
Previous research has sought to enhance the graph reasoning capabilities of LLMs by supervised fine-tuning on synthetic graph data. While these led to specialized LLMs better at so…