activity
20232026
collaborators

6 papers

cs.CV2026

EgoIntrospect: An Egocentric Dataset and Benchmark for User-Centric Internal State Reasoning

Zeyu Wang, Chang Liu, Eduardus Tjitrahardja +22

Despite extensive efforts on egocentric video datasets and benchmarks, understanding users' internal states, which is crucial for enabling seamless AI assistant experiences, remain…

cs.LG2026

LOFT: Low-Rank Orthogonal Fine-Tuning via Task-Aware Support Selection

Lanxin Zhao, Bamdev Mishra, Pratik Jawanpuria +4

Orthogonal parameter-efficient fine-tuning (PEFT) adapts pretrained weights through structure-preserving multiplicative transformations, but existing methods often conflate two dis…

cs.AI2026

CogniFold: Always-On Proactive Memory via Cognitive Folding

Suli Wang, Yiqun Duan, Yu Deng +6

Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genui…

cs.CV2025

VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models

Fufangchen Zhao, Liao Zhang, Daiqi Shi +5

We propose VideoPerceiver, a novel video multimodal large language model (VMLLM) that enhances fine-grained perception in video understanding, addressing VMLLMs' limited ability to…

cs.CV2025

MAGI-1: Autoregressive Video Generation at Scale

Sand. ai, Hansi Teng, Hongyu Jia +36

We present MAGI-1, a world model that generates videos by autoregressively predicting a sequence of video chunks, defined as fixed-length segments of consecutive frames. Trained to…

cs.CV2023

TransNeXt: Robust Foveal Visual Perception for Vision Transformers

Dai Shi

Due to the depth degradation effect in residual connections, many efficient Vision Transformers models that rely on stacking layers for information exchange often fail to form suff…