Technology

Spotify Posts Podcast Video Incident Report, Details June 24 Outage

Spotify's video transcoding infrastructure hit max capacity on June 24, delaying podcast publishes for hours. The postmortem names four converging causes, including a 10% resource utilization bug and a batch job that shouldn't have been running.

Spotify's engineering team has published an incident report on the June 24 podcast video outage, when transcoding infrastructure hit max capacity and delayed episode publishes for hours. The postmortem names four converging causes: insufficient headroom, a concurrent batch reprocessing job, a recent quality change that raised per-item cost, and a resource scheduling bug that left 10% of compute idle. Spotify has already increased transcoding capacity by 67%, fixed the scheduling bug, and committed to earlier alerting, better prioritization, and faster creator notification when publishing goes wrong.

What broke on June 24

Spotify's engineering team has published an incident report covering the June 24 publishing delay, the most significant of several reliability issues podcast creators have experienced over the past two months. The postmortem is unusually candid about both the failure and the response.

When a creator publishes a podcast episode, the audio and video content goes through transcoding and content analysis before it becomes available to users. On June 24, the video transcoding infrastructure reached maximum capacity. Episodes that would normally publish within minutes sat in a queue for hours. Some creators re-uploaded episodes that hadn't appeared, which added to the load. The system should have confirmed the upload was received and queued, and it didn't.

The four causes

Four factors converged. The transcoding infrastructure was running with insufficient headroom to absorb spikes in content delivery. A scheduled batch processing job, which reprocesses existing episodes to stay compatible with changes to the playback systems, was running at the same time and was consuming additional capacity. A recent change to the video pipeline, intended to deliver better quality at lower bitrates, had increased the per-item processing cost without the capacity plan accounting for it. A software bug from a recent infrastructure migration was underutilizing available compute by about 10%.

The combination of those four factors was enough to take the pipeline from "tight but functional" to "publishing queue for hours."

The timeline

The first alerts fired at 13:30 UTC, but they were not immediately recognized as a broader capacity issue. The video delivery spike pushed transcoding close to maximum at 15:00. The batch processing job was stopped at 16:35 to free capacity. Automated alerts confirmed the queue backlog exceeded thresholds at 17:34, and incident response began. Creator reports of missing episodes started coming in at 19:00. The software fix for the resource scheduling bug was deployed at 20:49. An additional processing cluster was brought online at 00:14 on June 25. All queues cleared by 01:02. Full confirmation that publishing was back to normal came at 07:30.

What Spotify changed

Three changes shipped immediately. Transcoding capacity is up about 67%, providing the headroom that was missing. The resource scheduling bug is fixed. Monitoring alerts fire earlier when capacity is approaching limits. The team has also formed a dedicated cross-team effort to invest in broader reliability work: better capacity planning that accounts for spikes and incident recovery, prioritization that puts real-time creator content ahead of background operations, and rate limiting and backpressure mechanisms throughout the pipeline.

The incident also exposed a communications gap. Many creators learned something was wrong from their audiences before they heard anything from Spotify. The team has committed to better processes and technical capabilities for early creator notification. For a platform where the publisher is the customer, the time-to-first-creator-message is itself a reliability metric, and the bar on it is now higher than it was before the outage.