4 papers
HarnessEval-W: Agentifying the Evaluation of Visual Worlds
Weiliang Chen, Haowen Sun, Jun Gao +40
A benchmark should deliver more than a scalar score: what makes an evaluation trustworthy is the reasoning that justifies the score. This is especially critical for world models, w…
Vision-Based Tactile Intelligence for Robotics: Sensing, Learning, and Embodied Manipulation
Peng Zhou, Jun Hu, Sihan Chen +12
Tactile sensing is essential for robots in contact-rich tasks, yet many tactile sensors still provide sparse, low-dimensional signals that do not capture sufficient information for…
Valhalla: A Layered Knowledge-State and Service-Governance Framework for Long-Term Scientific Knowledge Work
Yuyang Zheng, Nan Li, Wenxia Deng +3
As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-term memory have improved information ret…
Optimal Watermark Localization in Mixed-Source Large Language Model Texts
Jose H. Blanchet, T. Tony Cai, Xiang Li +3
Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evid…