4 citations · 7 across the 11 of their papers we have counts for
12 papers
Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification
Tianshu Zhang, Yan Wang, Ji Qi +1
Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language queries. While recent vision-language mod…
Do We Really Need External Tools to Mitigate Hallucinations? SIRA: Shared-Prefix Internal Reconstruction of Attribution
Tian Qin, Junzhe Chen, Yuqing Shi +3
Large vision-language models (LVLMs) often hallucinate when language priors dominate weak or ambiguous visual evidence. Existing contrastive decoding methods mitigate this problem…
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
V Team, Wenyi Hong, Xiaotao Gu +94
We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depen…
D3-Gym: Constructing Real-World Verifiable Environments for Data-Driven Discovery
Hanane Nour Moussa, Yifei Li, Zhuoyang Li +7
Despite recent progress in language models and agents for scientific data-driven discovery, advancing their capabilities is held back by the absence of verifiable environments repr…
SciNav: A General Agent Framework for Scientific Coding Tasks
Tianshu Zhang, Huan Sun
Autonomous science agents built on large language models (LLMs) are increasingly used to generate hypotheses, design experiments, and produce reports. However, prior work mainly ta…
EvoSchema: Towards Text-to-SQL Robustness Against Schema Evolution
Tianshu Zhang, Kun Qian, Siddhartha Sahai +4
Neural text-to-SQL models, which translate natural language questions (NLQs) into SQL queries given a database schema, have achieved remarkable performance. However, database schem…