INCIDENT رسمي مؤرشف

Incident with several GitHub Services

GitHub · آخر حالة مزوّد محفوظة: resolved

تقرير الخدمة الحالي ←

يُعرض نص المصدر الرسمي بلغته الأصلية.

هذا دليل تاريخي.

آخر حالة حدث محفوظة ليست حالة الخدمة الحالية. قد تكون قوائم المصدر ناقصة، واختفاء حدث لا يؤكد التعافي. افتح التقرير الحالي أو المصدر الرسمي لأدلة أحدث.

تفاصيل الحدث المحفوظة

حالة المزوّد
resolved
الأثر الذي يبلّغه المزوّد
critical
إنشاء سجل المزوّد
13 Sep 2026، 09:16:11 UTC
تحديث سجل المزوّد
15 Sep 2026، 21:47:16 UTC
بداية صريحة من المزوّد
13 Sep 2026، 09:16:11 UTC
نهاية صريحة من المزوّد
13 Sep 2026، 10:44:55 UTC

المكوّنات المتأثرة التي أبلغ عنها المزوّد

  • Actions br0l2tvcx85d
  • API Requests brv1bkgrwx7q
  • Pull Requests hhtssxt0f5v2
  • Issues kr09ddfgbfsf
  • Pages vg70hn9s2tyj

تصف هذه الارتباطات نطاق الحدث المبلّغ عنه. لا تثبت إتاحة المكوّن الحالية أو اعتماديات متحققة.

وقت إنشاء السجل ليس بالضرورة بداية العطل. تبقى الأوقات غير المبلّغ عنها غير متوفرة. لا نحسب مدة التعطل من أوقات الجمع.

تحديثات المزوّد في المراجعات المحفوظة

الأحدث أولًا. نعرض حتى 100 تحديث مختلف من آخر 20 مراجعة محتوى محفوظة. تُحفظ الصياغة المعدلة عند نفس وقت المزوّد منفصلة.

  1. resolved

    On September 13, 2026, between 08:43 and 10:44 UTC, GitHub experienced degraded availability across approximately 28 services, including Issues, Pull Requests, Actions, Codespaces, Pages, Notifications, Code Scanning, Git LFS, and new account signup. At peak, 8.8% of requests to create GitHub App installation access tokens failed. Token issuance for Actions workflows was also affected, impacting approximately 4% of workflows during the incident time frame. Creating issues through the web interface failed for about 96% of attempts, and signup failures were above 90%. <br /> <br />The cause was an internal data-cleanup job that began writing to a shared database cluster at 07:33 UTC. That cluster stores permission data read on nearly every authenticated request. The safeguard that was pacing the background job watched only one health signal — how far the database replicas were lagging — and that signal stayed low the whole time. It did not account for the load building on the primary itself, so the job kept writing while the primary quietly ran toward its limit. <br /><br />When the primary ran out of available connections, requests that needed it could not complete. First, there was no quick timeout on these database calls, so request handlers waited on the stalled database instead of failing fast, and the shared request-handling capacity degraded into site-wide errors. Second, a retry loop around token creation kept re-sending the writes that were already failing, which held the database saturated rather than letting it recover. <br /><br />Monitoring declared the incident at 08:50 UTC, but due to the broad impact and amplification from token creation, it took time to identify the source of the load. First responders mitigated by shedding internal load and pausing the job, and all services recovered by 10:44 UTC. <br /><br />To prevent recurrence, we are rate-limiting background jobs against shared, customer-serving databases by default, and adding automatic pausing and paging on primary-server load rather than replication lag alone. We are also surfacing running background work directly alongside database health signals so responders can see and pause it without leaving those dashboards, bounding retries in the token-issuing path, and adding request-level timeouts so one unhealthy database cannot consume shared web server capacity. In addition, we are breaking apart this database cluster to remove the single point of failure. We will be moving various service-specific data, including the authorization data, out of this shared cluster in the next two weeks.

    ظهر في مراجعة محفوظة في 06 Oct 2026، 13:58:39 UTC
  2. investigating

    Pull Requests is experiencing degraded performance. We are continuing to investigate.

    ظهر في مراجعة محفوظة في 06 Oct 2026، 13:58:39 UTC
  3. investigating

    We have reduced load on this cluster with internal load-shedding and are seeing signs of recovery but continue to monitor

    ظهر في مراجعة محفوظة في 06 Oct 2026، 13:58:39 UTC
  4. investigating

    We're seeing increased database replication delays on collab which is causing increased error rates in authorization endpoints and follow-on increased error rates across the system - we are investigating

    ظهر في مراجعة محفوظة في 06 Oct 2026، 13:58:39 UTC
  5. investigating

    Actions is experiencing degraded performance. We are continuing to investigate.

    ظهر في مراجعة محفوظة في 06 Oct 2026، 13:58:39 UTC
  6. investigating

    We are investigating reports of degraded availability for API Requests, Issues, Pages and Pull Requests

    ظهر في مراجعة محفوظة في 06 Oct 2026، 13:58:39 UTC