activity
20242026
collaborators

7 papers

cs.CV2026

Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR

Yulong Zhang, Tianyi Liang, Xinyue Huang +5

Optical Character Recognition (OCR) is fundamental to Vision-Language Models (VLMs) and high-quality data generation for LLM training. Yet, despite progress in average OCR accuracy…

cs.AI2025

Beyond Fixed Anchors: Precisely Erasing Concepts with Sibling Exclusive Counterparts

Tong Zhang, Ru Zhang, Jianyi Liu +2

Existing concept erasure methods for text-to-image diffusion models commonly rely on fixed anchor strategies, which often lead to critical issues such as concept re-emergence and e…

cs.AI2025

NCV: A Node-Wise Consistency Verification Approach for Low-Cost Structured Error Localization in LLM Reasoning

Yulong Zhang, Li Wang, Wei Du +7

Verifying multi-step reasoning in large language models is difficult due to imprecise error localization and high token costs. Existing methods either assess entire reasoning chain…

cs.LG2025

FedRS-Bench: Realistic Federated Learning Datasets and Benchmarks in Remote Sensing

Haodong Zhao, Peng Peng, Chiyu Chen +2

Remote sensing (RS) images are usually produced at an unprecedented scale, yet they are geographically and institutionally distributed, making centralized model training challengin…

cs.LG2025

InsightVision: A Comprehensive, Multi-Level Chinese-based Benchmark for Evaluating Implicit Visual Semantics in Large Vision Language Models

Xiaofei Yin, Yijie Hong, Ya Guo +4

In the evolving landscape of multimodal language models, understanding the nuanced meanings conveyed through visual cues - such as satire, insult, or critique - remains a significa…

cs.SD2025

U-GIFT: Uncertainty-Guided Firewall for Toxic Speech in Few-Shot Scenario

Jiaxin Song, Xinyu Wang, Yihao Wang +4

With the widespread use of social media, user-generated content has surged on online platforms. When such content includes hateful, abusive, offensive, or cyberbullying behavior, i…