Incident with Actions
GitHub · 마지막 저장 제공업체 상태: resolved
공식 출처 문구는 원어로 표시됩니다.
마지막 저장 사고 상태는 현재 서비스 상태가 아닙니다. 출처 목록은 불완전할 수 있고 소멸은 복구 확인이 아닙니다. 최신 증거는 현재 보고 또는 공식 출처를 확인하세요.
저장 사고 상세
- 제공업체 상태
- resolved
- 제공업체 영향
- critical
- 제공업체 기록 생성
- 26 Aug 2026, 15:11:58 UTC
- 제공업체 기록 갱신
- 27 Aug 2026, 02:24:56 UTC
- 제공업체가 명시한 시작
- 26 Aug 2026, 15:11:58 UTC
- 제공업체가 명시한 종료
- 26 Aug 2026, 18:01:30 UTC
제공업체가 보고한 영향받은 구성 요소
- Actions
br0l2tvcx85d - Pages
vg70hn9s2tyj
연결은 사고 보고 범위이며 현재 구성 요소 가용성이나 검증 의존 관계를 입증하지 않습니다.
기록 생성 시각이 반드시 장애 시작 시각인 것은 아닙니다. 보고되지 않은 시각은 알 수 없음으로 유지하며 수집 시각으로 중단 시간을 계산하지 않습니다.
저장 수정본의 제공업체 업데이트
최신순입니다. 최신 저장 내용 수정본 20개에서 최대 100개의 서로 다른 업데이트를 표시합니다. 같은 제공업체 시각의 문구 변경도 별도 보존됩니다.
- resolved
On August 26, 2026 from 15:02 to 15:45 UTC, Actions jobs failed to start. The following 2 hours until 17:40 UTC, Actions runs were delayed starting by more than 5 minutes as the system caught up with delayed load. This impact was triggered by saturation of writes to the database primary used by the service processing triggers for Actions workflows. The primary was failed over, but the system did not fully recover. The saturation was caused by growing daily peak load combined with an upstream issue in GitHub’s event processing infrastructure, https://www.githubstatus.com/incidents/hcbtzksccj2f, which caused burst amplification of already-high load. Downstream throttles that were later used to recover were set ~10% too high to protect the system. <br /><br />At 15:45 UTC, throttling combined with service restarts recovered the service’s core health. Those throttles were gradually raised between 15:54 and 17:22 to restore full webhook processing for Actions runs. This ramp was deliberately slow to ensure we did not re-overwhelm the system given our original throttling was now known to be incorrectly set. The queue of webhook events was fully burned down at 17:40 UTC. <br /><br />3.7% of larger-runner jobs, along with some scale-set self-hosted jobs, remained stuck in queued or “waiting for runner” state. We deployed a change to force-revoke jobs in this state, and they transitioned to failed at 18:40 UTC, about 50 minutes after incident mitigation. Releasing these jobs also freed hosted concurrency for larger-runner jobs. <br /><br />Customers using concurrency groups saw longer impact due to a separate issue where runners assigned to a subset of jobs disconnected before the force-revoke mitigation was deployed, which prevented runner acquisition from progressing and left jobs in a waiting-for-runner state. This was resolved at 01:00 UTC on August 27. <br /><br />Some runs triggered during the 15:02-15:45 UTC incident window encountered a bug that left them showing as queued even after service recovery. In the backend, these runs had already failed and will automatically move to canceled state 24 hours after creation. As follow-up, we are fixing the root cause of this queued state and improving our ability to bulk-cancel affected runs. <br /><br />Several changes to improve the general scalability of this part of Actions were already complete and deploying to production. Rollout of those changes will be complete within the next 24 hours. Further work to improve scale, resiliency, and more graceful degradation of Actions workflows are in flight. We are also taking a repair item to accelerate clearing of stuck queued or waiting jobs in similar future cases.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - monitoring
All inbound queues have recovered and Actions is operating as expected. 3.7% of jobs assigned to larger runners during the early stage of this incident are stuck waiting for runner assignment. Those will be canceled within the hour. Other runners are successfully processing all new jobs.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - monitoring
The degradation affecting Actions has been mitigated. We are monitoring to ensure stability.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
We are continuing to observe recovery and expect actions inbound queues to be back to normal in <30min. Work will continue to flow through the system subject to per-customer concurrency limits.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
We are continuing to observe recovery and delayed queues are burning down. Some customers will continue to see increased delays until all throttled work has been completed - we expect this within the next hour.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
Pages is operating normally.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
We believe we've identified and addressed the issue and are ramping traffic back up slowly to ensure it doesn't recur. Some customers will continue to see delays as we ramp up.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
We've identified an issue with a database primary and are failing over to a replica immediately
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
Pages is experiencing degraded performance. We are continuing to investigate.
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC - investigating
We are investigating reports of degraded availability for Actions
저장 수정본에서 확인 06 Oct 2026, 13:58:39 UTC