已歸檔的官方INCIDENT

Incident with several GitHub Services

GitHub · 最近儲存的提供方狀態:resolved

當前服務報告 →

官方來源文字以原始語言顯示。

這是歷史證據。

最近儲存的事件狀態不是當前服務狀態。來源列表可能不完整,事件從來源消失不確認恢復。請開啟當前報告或官方來源檢視較新證據。

已儲存事件詳情

提供方狀態
resolved
提供方影響程度
critical
提供方記錄建立時間
13 Sep 2026, 09:16:11 UTC
提供方記錄更新時間
15 Sep 2026, 21:47:16 UTC
提供方明確報告的開始時間
13 Sep 2026, 09:16:11 UTC
提供方明確報告的結束時間
13 Sep 2026, 10:44:55 UTC

提供方報告的受影響元件

  • Actions br0l2tvcx85d
  • API Requests brv1bkgrwx7q
  • Pull Requests hhtssxt0f5v2
  • Issues kr09ddfgbfsf
  • Pages vg70hn9s2tyj

這些關聯描述此事件的報告範圍。不能確認當前元件可用性或已核實依賴關係。

記錄建立時間未必是故障開始時間。未報告的時間保持不可用。我們不根據收集時間計算故障時長。

已儲存修訂中的提供方更新

最新優先。最多顯示最近 20 份已儲存內容修訂中的 100 條不同更新。同一提供方時間的文字修訂單獨保留。

  1. resolved

    On September 13, 2026, between 08:43 and 10:44 UTC, GitHub experienced degraded availability across approximately 28 services, including Issues, Pull Requests, Actions, Codespaces, Pages, Notifications, Code Scanning, Git LFS, and new account signup. At peak, 8.8% of requests to create GitHub App installation access tokens failed. Token issuance for Actions workflows was also affected, impacting approximately 4% of workflows during the incident time frame. Creating issues through the web interface failed for about 96% of attempts, and signup failures were above 90%. <br /> <br />The cause was an internal data-cleanup job that began writing to a shared database cluster at 07:33 UTC. That cluster stores permission data read on nearly every authenticated request. The safeguard that was pacing the background job watched only one health signal — how far the database replicas were lagging — and that signal stayed low the whole time. It did not account for the load building on the primary itself, so the job kept writing while the primary quietly ran toward its limit. <br /><br />When the primary ran out of available connections, requests that needed it could not complete. First, there was no quick timeout on these database calls, so request handlers waited on the stalled database instead of failing fast, and the shared request-handling capacity degraded into site-wide errors. Second, a retry loop around token creation kept re-sending the writes that were already failing, which held the database saturated rather than letting it recover. <br /><br />Monitoring declared the incident at 08:50 UTC, but due to the broad impact and amplification from token creation, it took time to identify the source of the load. First responders mitigated by shedding internal load and pausing the job, and all services recovered by 10:44 UTC. <br /><br />To prevent recurrence, we are rate-limiting background jobs against shared, customer-serving databases by default, and adding automatic pausing and paging on primary-server load rather than replication lag alone. We are also surfacing running background work directly alongside database health signals so responders can see and pause it without leaving those dashboards, bounding retries in the token-issuing path, and adding request-level timeouts so one unhealthy database cannot consume shared web server capacity. In addition, we are breaking apart this database cluster to remove the single point of failure. We will be moving various service-specific data, including the authorization data, out of this shared cluster in the next two weeks.

    在 06 Oct 2026, 13:58:39 UTC 的已儲存修訂中觀察到
  2. investigating

    Pull Requests is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已儲存修訂中觀察到
  3. investigating

    We have reduced load on this cluster with internal load-shedding and are seeing signs of recovery but continue to monitor

    在 06 Oct 2026, 13:58:39 UTC 的已儲存修訂中觀察到
  4. investigating

    We're seeing increased database replication delays on collab which is causing increased error rates in authorization endpoints and follow-on increased error rates across the system - we are investigating

    在 06 Oct 2026, 13:58:39 UTC 的已儲存修訂中觀察到
  5. investigating

    Actions is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已儲存修訂中觀察到
  6. investigating

    We are investigating reports of degraded availability for API Requests, Issues, Pages and Pull Requests

    在 06 Oct 2026, 13:58:39 UTC 的已儲存修訂中觀察到