Pablo Bernabeu Pérez
AI Safety Researcher
I am an AI Safety researcher focused on AI control, monitoring, and alignment evaluation. I am currently an independent Researcher backed by Coefficient Giving building benchmarks for distributed attacks in multi-agent coding environments.
Beyond my own research, I mentor early-career AI safety researchers. At Algoverse, I have supervised groups on trusted debate for AI control, cross-layer clustering for stochastic parameter decomposition, and model organisms of shutdown resistance. At SPAR, I co-mentor groups on ensemble monitoring and coordinated multi-agent sabotage. I also facilitated 4 groups of the Technical AI Safety Course at BlueDot Impact, mentored the first cohort of their AI Safety Project Sprint, and reviewed 20+ sprint projects.
Previously, I was an AI Safety Research Fellow at LASR Labs and MARS, and an external research collaborator at MATS. At LASR, I extended the AI Control framework to CoT monitoring, published at NeurIPS 2025. At MATS, I developed black-box sandbagging detection strategies, accepted at ICML 2026. At MARS, I evaluated capability-pressure RL as a recipe for training reward-hacking monitors.
Before transitioning to AI Safety, I was a Research Engineer at the Barcelona Supercomputing Center, where I developed SuSy (synthetic image detector), Aloe (state-of-the-art healthcare LLM), and CareQA (multilingual medical QA benchmark).
I hold an M.Sc. in Artificial Intelligence from the Polytechnic University of Catalonia and a B.Sc. in Computer Science from the Polytechnic University of Valencia. My master’s thesis was awarded the Best Applied AI Master’s Thesis Award by the Catalan Association of Artificial Intelligence.
selected publications
-
- In European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases, 2025