works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
collaborators

12 papers

cs.AI2026

NormAct: Benchmarking Embodied Agents' Proactive Compliance with Unspoken Social Norms

Shiyun Zhao, Xinwei Song, Tianyu Guo +7

The paper presents NormAct, a benchmark for evaluating whether embodied planners using multimodal large language models can infer and follow hidden social norms while completing ta…

cs.CL2026

DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning

Hengyu Fu, Tianyu Guo, Zixuan Wang +5

Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought (CoT). However, many questions require t…

cs.CL2026

Scaling Reasoning Tokens via RL and Parallel Thinking: Evidence From Competitive Programming

Qianfan Zhang, Tianyu Guo, Xuandi Ren +4

We study how to scale reasoning token budgets for competitive programming through two complementary approaches: training-time reinforcement learning (RL) and test-time parallel thi…

cs.CL2026

Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers

Yixiao Huang, Hanlin Zhu, Tianyu Guo +5

Large language models (LLMs) can acquire new knowledge through fine-tuning, but this process exhibits a puzzling duality: models can generalize remarkably from new facts, yet are a…

cs.AI2025

TongSIM: A General Platform for Simulating Intelligent Machines

Zhe Sun, Kunlun Wu, Chuanjian Fu +24

As artificial intelligence (AI) rapidly advances, especially in multimodal large language models (MLLMs), research focus is shifting from single-modality text processing to the mor…

cs.CV2025

SPEED-Q: Staged Processing with Enhanced Distillation towards Efficient Low-bit On-device VLM Quantization

Tianyu Guo, Shanwei Zhao, Shiai Zhu +1

Deploying Vision-Language Models (VLMs) on edge devices (e.g., smartphones and robots) is crucial for enabling low-latency and privacy-preserving intelligent applications. Given th…