activity
20242026
collaborators

6 papers

cs.AI2026

PetroBench: A Benchmark for Large Language Models in Petroleum Engineering

Xiang Wang, Tingting Zhang, Sen Wang +4

Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study develops a benchmark for LLMs…

cs.CV2026

FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling

Guixian Xu, Yide Liang, Zeli Su +5

Vision-language models have progressed rapidly, but Tibetan remains a severely underserved low-resource language due to the lack of reproducible training and evaluation infrastruct…

cs.IR2025

From Events to Trending: A Multi-Stage Hotspots Detection Method Based on Generative Query Indexing

Kaichun Wang, Yanguang Chen, Ting Zhang +7

LLM-based conversational systems have become a popular gateway for information access, yet most existing chatbots struggle to handle news-related trending queries effectively. To i…

cs.CL2025

CMHG: A Dataset and Benchmark for Headline Generation of Minority Languages in China

Guixian Xu, Zeli Su, Ziyin Zhang +4

Minority languages in China, such as Tibetan, Uyghur, and Traditional Mongolian, face significant challenges due to their unique writing systems, which differ from international st…

cs.CL2025

Multilingual Encoder Knows more than You Realize: Shared Weights Pretraining for Extremely Low-Resource Languages

Zeli Su, Ziyin Zhang, Guixian Xu +4

While multilingual language models like XLM-R have advanced multilingualism in NLP, they still perform poorly in extremely low-resource languages. This situation is exacerbated by…

cs.CL2024

DeTeCtive: Detecting AI-generated Text via Multi-Level Contrastive Learning

Xun Guo, Shan Zhang, Yongxin He +4

Current techniques for detecting AI-generated text are largely confined to manual feature crafting and supervised binary classification paradigms. These methodologies typically lea…