activity
20242026
collaborators

8 papers

cs.CV2026

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models

Martin Q. Ma, Willis Guo, Aditya Agrawal +4

Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and effic…

cs.CV2026

Act2See: Emergent Active Visual Perception for Video Reasoning

Martin Q. Ma, Yuxiao Qu, Aditya Agrawal +4

Vision-Language Models (VLMs) typically rely on static initial frames for video reasoning, restricting their ability to incorporate essential dynamic information as the reasoning p…

q-bio.NC2026

CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness

Haofei Yu, Yining Zhao, Lenore Blum +2

Despite remarkable advances, today's AI systems remain narrow in scope, falling short of the flexible, adaptive, and multisensory intelligence that characterizes human capabilities…

cs.AI2025

ORION: Teaching Language Models to Reason Efficiently in the Language of Thought

Kumar Tanmay, Kriti Aggarwal, Paul Pu Liang +1

Large Reasoning Models (LRMs) achieve strong performance in mathematics, code generation, and task planning, but their reliance on long chains of verbose "thinking" tokens leads to…

cs.CL2025

Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation

Keane Ong, Rui Mao, Deeksha Varshney +3

Counterfactual reasoning typically involves considering alternatives to actual events. While often applied to understand past events, a distinct form-forward counterfactual reasoni…

cs.CL2025

Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning

Jaedong Hwang, Kumar Tanmay, Seok-Jin Lee +5

Large Language Models (LLMs) have achieved strong performance in domains like mathematics, factual question answering, and code generation, yet their ability to reason on these tas…