ARCHIVED OFFICIAL INCIDENT

Actions delays in starting runs

GitHub · Last saved provider status: resolved

Current service report →
This is historical evidence.

The last saved event status is not the current service status. Source lists can be incomplete, and disappearance from a source does not confirm recovery. Open the current report or the official source for newer evidence.

Saved event details

Provider status
resolved
Provider impact
minor
Provider record created
24 Aug 2026, 13:56:54 UTC
Provider record updated
25 Aug 2026, 01:36:32 UTC
Explicit provider start
24 Aug 2026, 13:56:54 UTC
Explicit provider end
24 Aug 2026, 14:34:42 UTC

Affected components reported by the provider

  • Actions br0l2tvcx85d

These associations describe this event’s reported scope. They do not establish current component availability or verified dependencies.

A record creation time is not necessarily an outage start. Unreported times remain unavailable. We do not calculate downtime from collection times.

Provider updates in saved revisions

Latest first. We show up to 100 distinct updates from the latest 20 saved content revisions. Revised wording at the same provider time is retained separately.

  1. resolved

    On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. <br /> <br />The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. <br /><br />To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.

    Seen in a saved revision at 06 Oct 2026, 13:58:39 UTC
  2. monitoring

    The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

    Seen in a saved revision at 06 Oct 2026, 13:58:39 UTC
  3. investigating

    Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.

    Seen in a saved revision at 06 Oct 2026, 13:58:39 UTC
  4. investigating

    We are investigating reports of degraded performance for Actions

    Seen in a saved revision at 06 Oct 2026, 13:58:39 UTC