ARCHIVIERTER OFFIZIELLER INCIDENT

Actions delays in starting runs

GitHub · Letzter gespeicherter Anbieterstatus: resolved

Aktueller Dienstbericht →

Offizielle Quellentexte werden in Originalsprache angezeigt.

Historische Belege.

Letzter Ereignisstatus ist nicht aktueller Dienststatus. Quellen können unvollständig sein; Verschwinden bestätigt keine Erholung. Für neuere Belege aktuellen Bericht/Quelle öffnen.

Gespeicherte Ereignisdetails

Anbieterstatus
resolved
Anbieterauswirkungen
minor
Anbieterdatensatz erstellt
24 Aug 2026, 13:56:54 UTC
Anbieterdatensatz aktualisiert
25 Aug 2026, 01:36:32 UTC
Ausdrücklicher Anbieterbeginn
24 Aug 2026, 13:56:54 UTC
Ausdrückliches Anbieterende
24 Aug 2026, 14:34:42 UTC

Vom Anbieter gemeldete betroffene Komponenten

  • Actions br0l2tvcx85d

Zuordnungen beschreiben gemeldeten Ereignisumfang, keine aktuelle Komponentenverfügbarkeit oder verifizierten Abhängigkeiten.

Die Erstellung eines Datensatzes ist nicht zwingend der Ausfallbeginn. Nicht gemeldete Zeiten bleiben unbekannt. Aus Erfassungszeiten berechnen wir keine Ausfalldauer.

Anbieterupdates in gespeicherten Revisionen

Neueste zuerst. Bis zu 100 verschiedene Updates aus den letzten 20 gespeicherten Revisionen. Geänderte Formulierungen bei gleicher Anbieterzeit bleiben separat.

  1. resolved

    On August 24, 2026, between 13:33 UTC and 14:04 UTC, 3.8% of Actions runs experienced start delays over 5 minutes with 1.25% of Actions runs failing outright. <br /> <br />The incident was caused by a disk failure on a node hosting one of many service instances responsible for processing runner assignment events. Typically, pods on unhealthy nodes are removed and replaced automatically without impact. In this case, although the node was severely degraded and unable to perform disk operations, it continued sending healthy signals, preventing the system from immediately moving its work elsewhere. During this period, events assigned to the affected component accumulated until an automatic rebalance redirected processing to healthy components at 13:54 UTC. The queue backlog was cleared at 14:00 UTC, and processing returned to normal by 14:04 UTC. <br /><br />To prevent a recurrence, we are improving detection and automated remediation for unhealthy nodes that aren’t fully offline. We are also strengthening application-level resiliency, so stalled consumers are automatically removed quickly and their work reassigned without waiting for the affected node to recover.

    In gespeicherter Revision gesehen um 06 Oct 2026, 13:58:39 UTC
  2. monitoring

    The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.

    In gespeicherter Revision gesehen um 06 Oct 2026, 13:58:39 UTC
  3. investigating

    Failures while queuing and running Actions jobs for a subset of customers are now resolving. We are monitoring for full recovery.

    In gespeicherter Revision gesehen um 06 Oct 2026, 13:58:39 UTC
  4. investigating

    We are investigating reports of degraded performance for Actions

    In gespeicherter Revision gesehen um 06 Oct 2026, 13:58:39 UTC