16 papers
Semantic Intelligence Against CSAM: The PreventCSA@EU Ontology Framework for Classification and Investigation
Elias Tzortzakakis, Emmanouela Kokolaki, Evangelia Daskalaki +1
This work presents the PreventCSA@EU ontology, a semantically grounded framework designed to support the identification, classification, annotation, and analysis of online Child Se…
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
JunJia Guo, Yuhang Yao, Jiawei +2
We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based agents. Unlike prior code generation benchma…
Looped World Models
Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28
Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deeper models are expensive to deploy and prone to compounding error…
Not All Skills Help: Measuring and Repairing Agent Knowledge
Yixuan Wang, Yiyang Zhou, Yiming Liang +4
LLM agents can improve without weight updates by accumulating natural-language skills from experience, but current systems entrust every decision about which skills to keep and how…
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration
Jiaqi Liu, Shi Qiu, Mairui Li +33
Automating scientific discovery requires more than generating papers from ideas. Real research is iterative: hypotheses are challenged from multiple perspectives, experiments fail…
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?
Zixuan Lan, Luzhe Sun, Matthew R. Walter +1
Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to what extent such scores truly re…