From the 1 of 13 linked papers with an AI index.
13 papers
Implicit Reasoning Steering via Concept Chaining
Xiao Ye, Sanika Chavan, Yuxi Huang +4
The paper introduces Concept Chaining, a method that creates short natural-language paragraphs linking question entities to a target answer via intermediate concepts, and uses cont…
ASCII Art Turns LLMs into VLA Controllers
Yitao Jiang, Roy Xing, Luyang Zhao +3
Vision--Language--Action (VLA) controllers are often built by extending vision--language models (VLMs) with action supervision, relying on multimodal backbones with large data and…
FlowMo-WM: A World Model with Object Momentum and Hidden Ambient Drift
Yitao Jiang, Luyang Zhao, Muhao Chen +1
World models in robot learning predict future states from visual observations and actions, enabling agents to reason about the consequences of their controls. However, many action-…
VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
Zhaonan Li, Kyle R. Chickering, Bangzheng Li +13
A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipulate concept-level properties…
AQuaUI: Visual Token Reduction for GUI Agents with Adaptive Quadtrees
Yuankai Li, Tinghui Zhu, Ha Min Son +3
Large Multimodal Models (LMMs) have recently emerged as promising backbones for GUI-agent models, where high-resolution GUI screenshots are introduced to the prompts at each iterat…
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
Shahriar Kabir Nahin, Hadi Askari, Muhao Chen +1
The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deployment. While most videos can b…