4 papers
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
Siyoon Jin, Seongchan Kim, Dahyun Chung +5
Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internall…
Next-Depth Lookahead Tree
Jaeho Lee, Kangjin Kim, Gyeong Taek Lee
This paper proposes the Next-Depth Lookahead Tree (NDLT), a single-tree model designed to improve performance by evaluating node splits not only at the node being optimized but als…
Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
Junsu Kim, Naeun Kim, Jaeho Lee +3
The reasoning-based pose estimation (RPE) benchmark has emerged as a widely adopted evaluation standard for pose-aware multimodal large language models (MLLMs). Despite its signifi…
AssertBench: A Benchmark for Evaluating Self-Assertion in Large Language Models
Jaeho Lee, Atharv Chowdhary
Recent benchmarks have probed factual consistency and rhetorical robustness in Large Language Models (LLMs). However, a knowledge gap exists regarding how directional framing of fa…