works on

From the 1 of 25 linked papers with an AI index.

activity
20242026
collaborators

25 papers

cs.AI2026

Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

Taolin Han, Yuchen Zhang, Jinghang Wang +22

Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce S…

cs.CL2026

EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

Chengxi Liao, Tao Xu, Zulong Chen +7

The paper presents EntSQL, a benchmark designed to evaluate text-to-SQL models on enterprise tasks that require grounding in long, proprietary business documents, using a bilingual…

cs.CV2026

BabyVision: Visual Reasoning Beyond Language

Liang Chen, Weichu Xie, Yiyan Liang +27

While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile…

cs.LG2026

More Than Memory: Task-Conditioned Signed FFN Writes in Long-Context Retrieval

Zhibo Yang

FFNs are often treated as parametric memories. In long-context retrieval, however, the sharper question is not only what they store, but whether their native residual writes push t…

cs.CV2026

CodePercept: Code-Grounded Visual STEM Perception for MLLMs

Tongkun Guan, Zhibo Yang, Jianqiang Wan +10

When MLLMs fail at Science, Technology, Engineering, and Mathematics (STEM) visual reasoning, a fundamental question arises: is it due to perceptual deficiencies or reasoning limit…

cs.RO2026

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Qiuyue Wang, Mingsheng Li, Jian Guan +37

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generali…