3 papers
cs.AI2025
LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?
Omar Choukrani, Idriss Malek, Daniil Orel +5
When an interactive benchmark reports a single success rate for a language-model agent, it is rarely clear what that number measures. A failure can come from perception, ambiguous…
cs.LG2025
Loss-Guided Auxiliary Agents for Overcoming Mode Collapse in GFlowNets
Idriss Malek, Aya Laajil, Abhijith Sharma +2
Although Generative Flow Networks (GFlowNets) are designed to capture multiple modes of a reward function, they often suffer from mode collapse in practice, getting trapped in earl…
cs.LG2024
Free Lunch in the Forest: Functionally-Identical Pruning of Boosted Tree Ensembles
Youssouf Emine, Alexandre Forel, Idriss Malek +1
Tree ensembles, including boosting methods, are highly effective and widely used for tabular data. However, large ensembles lack interpretability and require longer inference times…