activity
20162026
most citedBERT & Family Eat Word Salad: Experiments with Text Understanding

24 citations · 36 across the 29 of their papers we have counts for

collaborators
Showing 2026Show all

6 papers · 1 filter

cs.CL2026

VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models

Zhiqi Huang, Vivek Datla, Zhichao Xu +3

Neural ranking models have become core components of modern information retrieval systems and important building blocks of AI systems such as retrieval-augmented generation (RAG) p…

cs.CL2026

The Machine's Internal Clock: Do LLMs Share Human Temporal Illusions?

Catherine Bao, Vivek Srikumar

Human perception of time is subjective. Well-documented temporal illusions show that the brain relies on context and relational cues for judging duration instead of tracking elapse…

cs.CL2026

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

Yu Wang, Shengyao Zhuang, Xueguang Ma +4

A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the right trade-off changes with th…

cs.CL2026

Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion

Maitrey Mehta, Nishant Subramani, Zhichao Xu +2

All languages are equal; when it comes to tokenization, some are more equal than others. Tokens are the hidden currency that dictate the cost and latency of access to contemporary…

cs.CL2026

InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis

Oliver Bentham, Vivek Srikumar

Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging. Benchmarks derived from published studi…

cs.IR2026

LACONIC: Dense-Level Effectiveness for Scalable Sparse Retrieval via a Two-Phase Training Curriculum

Zhichao Xu, Shengyao Zhuang, Crystina Zhang +5

While dense retrieval models have been the standard for state-of-the-art information retrieval, their deployment is often constrained by high memory requirements and reliance on GP…