collaborators

6 papers

eess.IV2026

SpaCE: Rethinking Spatial Capacity and Generalization in Multi-Frame Multimodal Large Language Models

Mariana Costa, Camila Ferreira, Alberlucia Rafael Soarez +1

Multi-modal large language models (MLLMs) have achieved remarkable empirical progress in spatial understanding through large-scale training on spatial visual question answering dat…

stat.ML2026

Demystifying Low-Rank Knowledge Distillation in Large Language Models: Convergence, Generalization, and Information-Theoretic Guarantees

Alberlucia Rafael Soarez, Daniel Kim, Mariana Costa +1

Knowledge distillation has emerged as a powerful technique for compressing large language models (LLMs) into efficient, deployable architectures while preserving their advanced cap…

cs.CY2026

The Ghost in the Grammar: Methodological Anthropomorphism in AI Safety Evaluations

Mariana Lins Costa

This essay offers a philosophical analysis of the field of AI safety based on recent technical reports, with particular focus on Anthropic's study on "agentic misalignment" in fron…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.CL2026

Enhancing Self-Correction in Large Language Models through Multi-Perspective Reflection

Mariana Costa, Alberlucia Rafael Soarez, Daniel Kim +1

While Chain-of-Thought (CoT) prompting advances LLM reasoning, challenges persist in consistency, accuracy, and self-correction, especially for complex or ethically sensitive tasks…

cs.AI2025

"They parted illusions -- they parted disclaim marinade": Misalignment as structural fidelity in LLMs

Mariana Lins Costa

The prevailing technical literature in AI Safety interprets scheming and sandbagging behaviors in large language models (LLMs) as indicators of deceptive agency or hidden objective…