AI

Google DeepMind AGI Safety Team Summarizes Recent Work

Google DeepMind's AGI Safety and Alignment Team (ASAT) published a July 2026 summary covering chain-of-thought monitoring, control, deep alignment, and interpretability.

Google DeepMind's AGI Safety and Alignment Team published a July 2026 summary covering progress in chain-of-thought monitoring, AI control, deep alignment, interpretability, and the Frontier Safety Framework. The team strengthened the FSF with misalignment considerations and continues work on risk assessments for frontier models.

The AGI Safety and Alignment Team (ASAT) at Google DeepMind, the main group working directly on technical approaches to existential risk from AI systems, has published a summary of its recent work[reference:131][reference:132]. The update covers progress across multiple research areas including chain-of-thought monitoring, AI control, deep alignment, and interpretability[reference:133].

Chain-of-Thought Monitoring

The team made significant progress on chain-of-thought monitorability. While the prevailing opinion was that chain-of-thought is unfaithful and effectively useless, the team argued this doesn't apply to difficult tasks requiring significant reasoning. They introduced "necessity" as a theoretically principled metric to quantify the extent to which the necessity argument applies for new model architectures[reference:134].

AI Control

Preparing for scenarios where chain-of-thought monitorability is lost, the team published initial thinking on AI control in collaboration with UK AISI. They worked with Google's security teams to build out monitoring systems and integrate them into internally deployed agents[reference:135].

Deep Alignment

The team shifted approach from conceptual alignment to working on current models. They argued that current models are likely similar enough to future AI systems that can substantially accelerate AI research, so progress on aligning current models has a good chance of directly transferring[reference:136].

Interpretability

After finding sparse autoencoders (SAEs) limited for downstream tasks, the team pivoted to other topics. They focus on areas that depend on the comparative advantages of interpretability researchers. Work includes probes, production misuse mitigations, and tooling for internal and external research[reference:137].

Amplified Oversight

The team studied obfuscated arguments in which a dishonest debater decomposes an easy problem into many hard subproblems. They published work making progress on obfuscated arguments by allowing the opposing debater to provide probability estimates over claims[reference:138].

Frontier Safety Framework

The team substantially strengthened the Frontier Safety Framework (FSF) and became the first company to introduce a section on misalignment[reference:139]. The update added clarity on security and deployment mitigations for various levels of dangerous capabilities and introduced Tracked Capability Levels (TCLs) for less severe risks[reference:140].

Risk Assessments

The team conducts risk assessments in collaboration with partner teams. An example is the Gemini 3.5 Pro risk assessment, drawing on chain-of-thought legibility research and other work from ASAT[reference:141].

Broader Impact

The summary reflects the team's focus on technical AGI safety and alignment as AI capabilities continue to accelerate. The work spans monitoring, control, alignment, and governance, with an emphasis on preparing for both near-term and existential risks from advanced AI systems.