4 papers
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
Yujin Zhou, Mingxuan Zheng, Chuxue Cao +4
As large language models are increasingly deployed as tool-augmented legal agents, they introduce agentic hallucinations where tool-call and reasoning errors cascade into fabricate…
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Zhixiang Liang, Yifei Liu, Yidan Huang +5
Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long…
Joint Discrete-Continuous Flow Matching for Open-Vocabulary Inverse Design of Multilayer Optical Coatings
Zhiyi Li, Yuheng Jin, Yidan Huang +3
Amortized neural inverse design typically remains closed-world: component choices are fixed vocabulary tokens, coordinate grids are frozen at training time, and continuous variable…
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Songlin Yang, Haobin Zhong, Ruilin Zhang +23
The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community tr…