已归档的官方INCIDENT

Incident with GitHub.com

GitHub · 最近保存的提供方状态:resolved

当前服务报告 →

官方来源文字以原始语言显示。

这是历史证据。

最近保存的事件状态不是当前服务状态。来源列表可能不完整,事件从来源消失不确认恢复。请打开当前报告或官方来源查看较新证据。

已保存事件详情

提供方状态
resolved
提供方影响程度
critical
提供方记录创建时间
17 Aug 2026, 13:40:03 UTC
提供方记录更新时间
18 Aug 2026, 19:21:36 UTC
提供方明确报告的开始时间
17 Aug 2026, 13:40:03 UTC
提供方明确报告的结束时间
17 Aug 2026, 21:15:46 UTC

提供方报告的受影响组件

  • Webhooks 4230lsnqdsld
  • Git Operations 8l4ygp009s5s
  • Actions br0l2tvcx85d
  • API Requests brv1bkgrwx7q
  • Pull Requests hhtssxt0f5v2
  • Issues kr09ddfgbfsf
  • Copilot pjmpxvq2cmr2
  • Pages vg70hn9s2tyj

这些关联描述此事件的报告范围。不能确认当前组件可用性或已核实依赖关系。

记录创建时间未必是故障开始时间。未报告的时间保持不可用。我们不根据采集时间计算故障时长。

已保存修订中的提供方更新

最新优先。最多显示最近 20 份已保存内容修订中的 100 条不同更新。同一提供方时间的文字修订单独保留。

  1. resolved

    On August 17, 2026, from 13:28–21:15 UTC (7h 47m), GitHub.com experienced elevated errors and latency across Issues, Pull Requests, APIs, Actions, and Copilot. At peak, web/API error rates were approximately 20%, while archive and raw-content downloads reached approximately 50%. SAML/OIDC authentication, SCIM, and Team Sync were also affected, as well as Actions workflows in GHEC with Data Residency that depend on public workflow step definitions hosted on GitHub.com. Most services recovered by 16:36 UTC as our Central US datacenter recovered; Actions was degraded until approximately 18:03 UTC; and Copilot Token Service fully recovered by 21:02. <br /><br />Some of the failing traffic was moved from Central US to Northern Virginia where it was served successfully until the network failure in Central US was debugged and resolved. Delayed replies to a single internal endpoint triggered a latent retry bug in VS Code that amplified traffic by approximately 10x and caused delayed recovery for the Copilot Token Service. <br /><br />The immediate cause of the failure was network saturation on load balancers in Central US due to a new peak in traffic. Originally this was caused by an Istio sidecar pod reaching its concurrency limits and failing to auto scale correctly because of a misconfigured policy that watched host service but not sidecar limits. One failure cascaded to more and eventually four HAProxy nodes exhausted their flow limits, degrading the gateway auth path and causing widespread authentication latency and failures. The problem was worsened by optimistic retry logic which overloaded internal load balancers. Pausing HAProxy on those nodes simultaneously produced immediate broad recovery. <br /><br />The retry storm in Northern VA was fixed by 1) temporarily reducing gateway retry logic with a PR and 2) blocking inbound Copilot Token Service token requests at the load balancers with a 403, and then gradually ramping back up traffic per-site to allow callers to succeed. <br /><br />Residual Copilot authentication failures continued because client retry behavior amplified load: a failed token operation could generate many extra requests and enter a retry loop. Copilot Token Service traffic increased from a normal 7–9K RPS to 70–100K RPS. Reducing gateway authentication retries and blocking retry-triggering responses stabilized Copilot Token Service and completed recovery. <br /><br />Complicating factors that impeded recovery included a number of scraping attacks on codeload endpoints. <br /><br />To prevent recurrence, our follow-up actions include: <br /><br />- Correcting autoscaling policies to account for service-mesh sidecar concurrency and capacity. <br /><br />- Auditing Istio request, concurrency, and scaling limits across affected services. <br /><br />- Reviewing retry limits and backoff behavior across gateways and clients. <br /><br />- Addressing the VS Code retry behavior that amplified Copilot token traffic. <br /><br />- Improving load-balancer capacity monitoring and regional failover safeguards.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  2. investigating

    We are continuing to apply mitigations to address sporadic Copilot authentication failures in some applications. We expect full recovery within the next 30 minutes. Copilot usage via the GitHub CLI and GitHub App are unaffected.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  3. investigating

    Issues is operating normally.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  4. investigating

    We are continuing to investigate sporadic failures affecting Copilot authentication in some applications. Copilot usage via the GitHub CLI and GitHub App are unaffected.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  5. investigating

    We are continuing to investigate sporadic authentication failures. We have partially disabled authentication token retries and have seen improvement, and we are monitoring impact before fully applying this mitigation.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  6. investigating

    API Requests is operating normally.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  7. investigating

    API Requests is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  8. investigating

    The degradation affecting Git Operations has been mitigated. We are monitoring to ensure stability.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  9. investigating

    We identified the problematic component and have taken corrective actions, but we are seeing residual impact in the form of sporadic authentication failures. We are continuing to apply additional mitigations and investigate the remaining impact.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  10. investigating

    Issues is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  11. investigating

    We identified the problematic component and have taken corrective actions, but we are seeing residual impact across numerous services. We are continuing to apply additional mitigations and investigate the remaining impact.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  12. investigating

    Git Operations is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  13. investigating

    The degradation affecting API Requests, Actions, Git Operations, Issues, Pages, Pull Requests and Webhooks has been mitigated. We are monitoring to ensure stability.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  14. investigating

    We identified the problematic component and have taken corrective actions. There are strong signs of recovery but we are still working to completely restore service, with error rates still remaining slightly elevated. We will post further updates as recovery continues.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  15. investigating

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are still working to identify the root cause and will continue to post updates as we learn more and perform mitigation.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  16. investigating

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations and will post updates as we progress.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  17. investigating

    Webhooks is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  18. investigating

    Git Operations is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  19. investigating

    Pages is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  20. investigating

    API Requests is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  21. investigating

    Webhooks is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  22. investigating

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. We are currently performing mitigations based on our investigation thus far and are monitoring for improvement.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  23. investigating

    Actions is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  24. investigating

    Pull Requests is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  25. investigating

    Issues is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  26. investigating

    Pull Requests is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  27. investigating

    Copilot is experiencing degraded availability. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  28. investigating

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. SAML and OIDC authentication, SCIM, and Team Sync are also impacted. Investigations are on-going and we will continue to provide updates as we discover more information.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  29. investigating

    We are experiencing high error rates around 20% for web experiences and api traffic. Archive downloads and raw repository content downloads are experiencing an approximate 50% error rate. Investigations are on-going into the root cause, and updates will continue to be provided as we investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  30. investigating

    Pull Requests is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  31. investigating

    Issues is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  32. investigating

    We are seeing an approximate 20% error rate across numerous experiences including Pull Requests, Issues, and others. Investigations are currently under way and we will be posting updates as they become available

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  33. investigating

    Webhooks is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  34. investigating

    Actions is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  35. investigating

    API Requests is experiencing degraded performance. We are continuing to investigate.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到
  36. investigating

    We are investigating reports of impacted performance for some GitHub services.

    在 06 Oct 2026, 13:58:39 UTC 的已保存修订中观察到