collaborators

5 papers

cs.CL2026

S^2tory: Story Spine Distillation for Movie Script Summarization

Mingzhe Lu, Yanbing Liu, Qihao Wang +5

Movie scripts pose a fundamental challenge for automatic summarization due to their non-linear, cross-cut narrative structure, which makes surface-level saliency methods ineffectiv…

cs.CL2026

Evaluating Memory Capability in Continuous Lifelog Scenario

Jianjie Zheng, Zhichen Liu, Zhanyu Shen +6

Nowadays, wearable devices can continuously lifelog ambient conversations, creating substantial opportunities for memory systems. However, existing benchmarks primarily focus on on…

cs.CL2026

Think in Sentences: Explicit Sentence Boundaries Enhance Language Model's Capabilities

Zhichen Liu, Yongyuan Li, Yang Xu

Researchers have explored different ways to improve large language models (LLMs)' capabilities via dummy token insertion in contexts. However, existing works focus solely on the du…

cs.CV2025

DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search

Kartik Narayan, Yang Xu, Tian Cao +7

Multimodal Large Language Models (MLLMs) in real-world applications require access to external knowledge sources and must remain responsive to the dynamic and ever-changing real-wo…

cs.CV2025

VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety

Shruti Palaskar, Leon Gatys, Mona Abdelrahman +9

Safety evaluation of multimodal foundation models often treats vision and language inputs separately, missing risks from joint interpretation where benign content becomes harmful i…