posts
- ai-safety
- monitoring
- ai-control
- llms
•
•
•
-
Unfaithful Reasoning Can Fool Chain-of-Thought Monitoring
Can we catch scheming AI by reading its chain-of-thought? We stress-test CoT monitoring and find that unfaithful reasoning can bypass it.
-
Weak Control
An exploration of AI control approaches and their limitations.
-
Catching Up on the Weird World of LLMs
A tour through the strange and surprising behaviors of large language models.