2 citations · 2 across the 7 of their papers we have counts for
60 papers
UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs
Xuexiong Yin, Zechuan Chen, Yongsen Zheng +5
Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personal…
ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models
Yihao Wang, Zijian He, Jie Ren +1
Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different actions under different task contexts. How…
How Do Document Parsers Break? Auditing Structural Vulnerability in Document Intelligence
Yue Chen, Yihao Wang, Ziyi Tang +2
Document Layout Analysis (DLA) pipelines provide structured page representations for retrieval-augmented generation, long-document question answering, and other document intelligen…
Kolmogorov-Arnold Fourier Networks
Jusheng Zhang, Yijia Fan, Kaitong Cai +2
Although Kolmogorov-Arnold-based interpretable networks (KANs) possess strong theoretical expressiveness, they suffer from severe parameter explosion and limited ability to capture…
PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
Zeqing Wang, Keze Wang, Lei Zhang
Driven by the growing capacity and training scale, Text-to-Video (T2V) generation models have recently achieved substantial progress in video quality, length, and instruction-follo…
LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
Jinzhou Tang, Sidi Liu, Waikit Xiu +2
A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critica…