6 citations · 8 across the 3 of their papers we have counts for
4 papers · 1 filter
EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
Yunhan Wang, Jiaan Wang, Lianzhe Huang +2
Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks. Existing benchmarks such as BrowseComp rely…
RECALL: A Benchmark for LLMs Robustness against External Counterfactual Knowledge
Yi Liu, Lianzhe Huang, Shicheng Li +5
LLMs and AI chatbots have improved people's efficiency in various fields. However, the necessary knowledge for answering the question may be beyond the models' knowledge boundaries…
Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification
Zihan Wang, Peiyi Wang, Lianzhe Huang +2
Hierarchical text classification is a challenging subtask of multi-label classification due to its complex label hierarchy. Existing methods encode text and label hierarchy separat…
Explicit Interaction Network for Aspect Sentiment Triplet Extraction
Peiyi Wang, Tianyu Liu, Damai Dai +3
Aspect Sentiment Triplet Extraction (ASTE) aims to recognize targets, their sentiment polarities and opinions explaining the sentiment from a sentence. ASTE could be naturally divi…