17 papers
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Zhen Fang, Yu Zeng, Wenxuan Huang +17
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding couple…
ExpressionCueLens: A Cross-Cultural Analysis of Human-AI Companion Conversations on Social Media
Lynnette Hui Xian Ng, Yunze Xiao, Lionel Z. Wang +2
The paper presents ExpressionCueLens, a framework for categorizing anthropomorphic expressions in human‑AI companion conversations, and uses it to compare how Reddit and XiaoHongSh…
NüshuVoice: Reviving the Voice of Endangered Nüshu with Pitch-Aware Text-to-Speech
Hongkun Yang, Xinhui Yi, Xiyan Zhao +13
Nüshu is an endangered phonetic script historically used by women in Jiangyong County, southern Hunan, China. While existing computational studies of Nüshu mainly focus on textua…
Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs
Jianan Li, Simeng Qin, Xiaojun Jia +5
Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world applications. However, their e…
Contexting as Recommendation: Evolutionary Collaborative Filtering for Context Engineering
Jiachen Zhu, Zhuoying Ou, Congmin Zheng +9
Large Language Models (LLMs) are highly sensitive to their input contexts, motivating the development of automated context engineering. However, existing methods predominantly trea…
SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generation
Tianfei Ren, Zhipeng Yan, Yiming Zhao +13
While text-to-image models have made strong progress in visual fidelity, faithfully realizing complex visual intents remains challenging because many requirements must be tracked a…