5 papers
Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting
Jingyi Kang, Junyu Lu, Bo Xu +4
Large language models (LLMs) require robust toxicity evaluation beyond explicit wording. This setting remains underexplored in Chinese, where toxicity may combine semantic indirect…
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination
Qiyao Wang, Xinyi Chen, Longze Chen +4
Patent examination is a complex, multi-stage process requiring both technical expertise and legal reasoning, increasingly challenged by rising application volumes. Prior benchmarks…
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation?
Qiyao Wang, Haoran Hu, Longze Chen +4
With the advancement of multimodal large language models (MLLMs) and coding agents, the website development has shifted from manual programming to agent-based project-level code sy…
Propagating Similarity, Mitigating Uncertainty: Similarity Propagation-enhanced Uncertainty for Multimodal Recommendation
Xinzhuo Wu, Hongbo Wang, Yuan Lin +3
Multimodal Recommendation (MMR) systems are crucial for modern platforms but are often hampered by inherent noise and uncertainty in modal features, such as blurry images, diverse…
IPBench: Benchmarking the Knowledge of Large Language Models in Intellectual Property
Qiyao Wang, Guhong Chen, Hongbo Wang +20
Intellectual Property (IP) is a highly specialized domain that integrates technical and legal knowledge, making it inherently complex and knowledge-intensive. Recent advancements i…