activity
20242026
collaborators

8 papers

cs.DB2026

Semantic Intelligence Against CSAM: The PreventCSA@EU Ontology Framework for Classification and Investigation

Elias Tzortzakakis, Emmanouela Kokolaki, Evangelia Daskalaki +1

This work presents the PreventCSA@EU ontology, a semantically grounded framework designed to support the identification, classification, annotation, and analysis of online Child Se…

cs.CL2026

On the Predictive Power of Representation Dispersion in Language Models

Yanhong Li, Ming Li, Karen Livescu +1

We show that a language model's ability to predict text is tightly linked to the breadth of its embedding space: models that spread their contextual representations more widely ten…

cs.LG2026

Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting

Xinghong Fu, Yanhong Li, Georgios Papaioannou +1

Learning time series foundation models has been shown to be a promising approach for zero-shot time series forecasting across diverse time series domains. Insofar as scaling has be…

cs.CL2025

Distilling to Hybrid Attention Models via KL-Guided Layer Selection

Yanhong Li, Songlin Yang, Shawn Tan +4

Distilling pretrained softmax attention Transformers into more efficient hybrid architectures that interleave softmax and linear attention layers is a promising approach for improv…

cs.CL2025

OKBench: Democratizing LLM Evaluation with Fully Automated, On-Demand, Open Knowledge Benchmarking

Yanhong Li, Tianyang Xu, Kenan Tang +3

Knowledge-intensive question answering is central to large language models (LLMs) and is typically assessed using static benchmarks derived from sources like Wikipedia and textbook…

cs.CL2025

Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs

Yanhong Li, Zixuan Lan, Jiawei Zhou

Large language models (LLMs) and their multimodal variants can now process visual inputs, including images of text. This raises an intriguing question: can we compress textual inpu…