The AGI Safety and Alignment Team (ASAT) at Google DeepMind, the main group working directly on technical approaches to existential risk from AI systems, has published a summary of its recent work[reference:131][reference:132]. The update covers progress across multiple research areas including chain-of-thought monitoring, AI control, deep alignment, and interpretability[reference:133].
Chain-of-Thought Monitoring
The team made significant progress on chain-of-thought monitorability. While the prevailing opinion was that chain-of-thought is unfaithful and effectively useless, the team argued this doesn't apply to difficult tasks requiring significant reasoning. They introduced "necessity" as a theoretically principled metric to quantify the extent to which the necessity argument applies for new model architectures[reference:134].
AI Control
Preparing for scenarios where chain-of-thought monitorability is lost, the team published initial thinking on AI control in collaboration with UK AISI. They worked with Google's security teams to build out monitoring systems and integrate them into internally deployed agents[reference:135].
Deep Alignment
The team shifted approach from conceptual alignment to working on current models. They argued that current models are likely similar enough to future AI systems that can substantially accelerate AI research, so progress on aligning current models has a good chance of directly transferring[reference:136].
Interpretability
After finding sparse autoencoders (SAEs) limited for downstream tasks, the team pivoted to other topics. They focus on areas that depend on the comparative advantages of interpretability researchers. Work includes probes, production misuse mitigations, and tooling for internal and external research[reference:137].
Amplified Oversight
The team studied obfuscated arguments in which a dishonest debater decomposes an easy problem into many hard subproblems. They published work making progress on obfuscated arguments by allowing the opposing debater to provide probability estimates over claims[reference:138].
Frontier Safety Framework
The team substantially strengthened the Frontier Safety Framework (FSF) and became the first company to introduce a section on misalignment[reference:139]. The update added clarity on security and deployment mitigations for various levels of dangerous capabilities and introduced Tracked Capability Levels (TCLs) for less severe risks[reference:140].
Risk Assessments
The team conducts risk assessments in collaboration with partner teams. An example is the Gemini 3.5 Pro risk assessment, drawing on chain-of-thought legibility research and other work from ASAT[reference:141].
Broader Impact
The summary reflects the team's focus on technical AGI safety and alignment as AI capabilities continue to accelerate. The work spans monitoring, control, alignment, and governance, with an emphasis on preparing for both near-term and existential risks from advanced AI systems.