1 citations · 2 across the 11 of their papers we have counts for
11 papers
VeriScene: Reconstructing Crime Scenes from Legal Evidence via World-Model Agent
Kevin Chuanpu Fu, Yongsen Zheng, Zee Kin Yeong +1
World models take multimodal inputs like text, photos, and diagrams to generate dynamic scenes in accordance with the laws of physics, thus opening a compelling application: fusing…
UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs
Xuexiong Yin, Zechuan Chen, Yongsen Zheng +5
Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personal…
How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence
Yue Chen, Yihao Wang, Ziyi Tang +2
Document Layout Analysis (DLA) pipelines provide structured page representations for retrieval-augmented generation, long-document question answering, and related applications. Yet…
MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
Xuanye Zhang, Yongsen Zheng, Zhuqin Xu +5
LLM-driven agents are capable of selecting external tools to complete users' tasks. However, attackers could compromise such process, steering agents toward inappropriate/wrong too…
Process-of-Thought Reasoning for Videos
Jusheng Zhang, Kaitong Cai, Jian Wang +3
Video understanding requires not only recognizing visual content but also performing temporally grounded, multi-step reasoning over long and noisy observations. We propose Process-…
Spectral Gating Networks
Jusheng Zhang, Yijia Fan, Kaitong Cai +5
Gating mechanisms are ubiquitous, yet a complementary question in feed-forward networks remains under-explored: how to introduce frequency-rich expressivity without sacrificing sta…