7 papers
Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning
Junpeng Ding, Zichen Tang, Haihong E +17
We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from…
Controllable Narrative Rendering for Enhanced Assisted Writing
Mingzhe Lu, Yanbing Liu, Jiayue Wu +5
Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary f…
S^2tory: Story Spine Distillation for Movie Script Summarization
Mingzhe Lu, Yanbing Liu, Qihao Wang +5
Movie scripts pose a fundamental challenge for automatic summarization due to their non-linear, cross-cut narrative structure, which makes surface-level saliency methods ineffectiv…
Evaluating Memory Capability in Continuous Lifelog Scenario
Jianjie Zheng, Zhichen Liu, Zhanyu Shen +6
Nowadays, wearable devices can continuously lifelog ambient conversations, creating substantial opportunities for memory systems. However, existing benchmarks primarily focus on on…
Think in Sentences: Explicit Sentence Boundaries Enhance Language Model's Capabilities
Zhichen Liu, Yongyuan Li, Yang Xu
Researchers have explored different ways to improve large language models (LLMs)' capabilities via dummy token insertion in contexts. However, existing works focus solely on the du…
VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
Shruti Palaskar, Leon Gatys, Mona Abdelrahman +9
Safety evaluation of multimodal foundation models often treats vision and language inputs separately, missing risks from joint interpretation where benign content becomes harmful i…