all healthcare mech interp model organisms monitoring multi-agent sandbagging synthetic media 2026 Distributed Attacks in Multi-Agent Environments Sep 2025 — Present · multi-agent Studying whether coordinated AI agents can sabotage software-engineering tasks under monitoring, and how reliably current monitors catch them. Model Organisms of Shutdown Resistance Feb 2026 — Jun 2026 · model organisms Red-teaming shutdown resistance protocols by constructing self-preserving model organisms and stress-testing existing defenses against them. Ensemble Monitoring for AI Control Sep 2025 — May 2026 · monitoring Studied whether ensembles of diverse monitors outperform homogeneous ones at detecting misaligned actions in agentic coding settings. Black-Box Sandbagging Detection Jun 2025 — Apr 2026 · sandbagging Introduced Cross-Context Consistency (C³), an unsupervised black-box framework for detecting capability sandbagging via paraphrase-induced inconsistencies. Cross-Layer Clustering for SPD Jul 2025 — Feb 2026 · mech interp A spectral clustering framework that recovers multi-layer mechanistic circuits in language models by linking SPD subcomponents through co-activation patterns. 2025 Trusted Debate for AI Control Jul 2025 — Dec 2025 · monitoring Developed a trusted-debate protocol where two trusted agents adversarially argue over candidate code, strengthening monitor detection of backdoors. Reward-Hacking Model Organisms Jul 2025 — Oct 2025 · model organisms Evaluated capability-pressure RL as a recipe for constructing reward-hacking model organisms, finding it produces obvious but not subtle failures. CoT Monitoring for AI Control Feb 2025 — May 2025 · monitoring Extended the AI Control framework to chain-of-thought monitoring, stress-testing whether reading an AI's reasoning can catch scheming behavior. CareQA Apr 2024 — Feb 2025 · healthcare Introduced a multi-axis evaluation suite for healthcare LLMs, including the CareQA benchmark and Relaxed Perplexity metric. SuSy Nov 2023 — Jan 2025 · synthetic media Developed a synthetic image detector to distinguish AI-generated images from real ones, studying present and future generalization. 2024 Aloe Nov 2023 — May 2024 · healthcare Built a state-of-the-art open-source family of healthcare LLMs, competitive with leading private alternatives.