activity
20242026
most citedToward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation

1 citations · 1 across the 15 of their papers we have counts for

collaborators
Showing 2025Show all

15 papers · 1 filter

cs.CV2025

HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing

Zixuan Bian, Ruohan Ren, Yue Yang +1

3D scene generation plays a crucial role in gaming, artistic creation, virtual reality, and many other domains. However, current 3D scene design still relies heavily on extensive m…

cs.CV2025

DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions

Xiaoyu Lin, Aniket Ghorpade, Hansheng Zhu +9

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limit…

cs.CL2025

The Media Bias Detector: A Framework for Annotating and Analyzing the News at Scale

Samar Haider, Amir Tohidi, Jenny S. Wang +4

Mainstream news organizations shape public perception not only directly through the articles they publish but also through the choices they make about which topics to cover (or ign…

cs.LG2025

Probabilistic Soundness Guarantees in LLM Reasoning Chains

Weiqiu You, Anton Xue, Shreya Havaldar +4

In reasoning chains generated by large language models (LLMs), initial errors often propagate and undermine the reliability of the final conclusion. Current LLM-based error detecti…

cs.CL2025

Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap

Andrew Zhu, Chris Callison-Burch

Imagine AI assistants that enhance conversations without interrupting them: quietly providing relevant information during a medical consultation, seamlessly preparing materials as…

cs.AI2025

Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D

Artemis Panagopoulou, Le Xue, Honglu Zhou +6

Real-world decision-making often begins with identifying which modality contains the most relevant information for a given query. While recent multimodal models have made impressiv…