anticipatory planning 1benchmark evaluation 1computer-use agents 1error analysis 1gui automation 1latency reduction 1multimodal models 1policy trees 1reliability 1task scoring 1
From the 2 of 14 linked papers with an AI index.
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Assessment of Generative Named Entity Recognition in the Era of Large Language Models
Qi Zhan, Yile Wang, Hui Huang
Named entity recognition (NER) is evolving from a sequence labeling task into a generative paradigm with the rise of large language models (LLMs). We conduct a systematic evaluatio…
cs.CL2026
Disentangling Language Roles in Multilingual LLM Task Execution
Qishi Zhan, Minxuan Hu, Seoyeon Jang +7
Multilingual LLMs are increasingly used when instruction, source content, and required response languages do not coincide. Existing benchmarks have expanded multilingual instructio…
cs.CL2026
How Order-Sensitive Are LLMs? OrderProbe for Deterministic Structural Reconstruction
Yingjie He, Zhaolu Kang, Kehan Jiang +19
Large language models (LLMs) excel at semantic understanding, yet their ability to reconstruct internal structure from scrambled inputs remains underexplored. Sentence-level restor…