Delayed notifications
Jul 28, 20:57 UTC Resolved - Notification latency for the affected shard is now under 5s. Jul 28, 20:38 UTC Identified - We are investigating delays to build and job notifications for a subset of customers.
Jul 28, 20:57 UTC Resolved - Notification latency for the affected shard is now under 5s. Jul 28, 20:38 UTC Identified - We are investigating delays to build and job notifications for a subset of customers.
Jun 18, 09:43 UTC Resolved - We have seen a full recovery of services. Jun 18, 09:17 UTC Monitoring - A fix has been deployed, services are recovering. Jun 18, 08:59 UTC Update - We've identified the issue, we're continuing working on rolling out a fix. Impact is restricted to the builds create API for a subset of tenants. Jun 18, 08:05 UTC Identified - We're seeing increased latency and error rat…
Jun 12, 00:16 UTC Resolved - The mitigation applied before the last update had the intended effect, and we have seen recovery in REST API latency. Jun 11, 23:51 UTC Update - We've isolated the issue to elevated load on our REST API service and have mitigated the issue. The agent and stacks API isn’t affected. Jun 11, 23:49 UTC Monitoring - We've isolated the issue to elevated load on our REST API…
Jun 11, 00:51 UTC Resolved - Between 00:05 - 00:34 UTC, a subset of customers experienced increased latency and timeout errors on the Agent API. This impacts job assignment. At peak impact, we saw an error rate of 1.3% of requests and job acceptance latency up to 53s. Jun 11, 00:32 UTC Investigating - We're observing increased latency and error rates for a subset of our customers on the Agent API.…
May 30, 00:30 UTC Resolved - We have received reports email deliveries have not been working, affecting signup and invite emails as well as build notification emails.This issue has now been resolved.
May 28, 21:18 UTC Resolved - This incident has been resolved. May 28, 20:47 UTC Identified - We've spotted that something has gone wrong. We're currently investigating the issue, and will provide an update soon. May 28, 20:20 UTC Investigating - We are investigating delays to build and job notifications for a subset of customers.
May 26, 10:38 UTC Resolved - We think the impact from the issue is over. May 26, 10:37 UTC Update - We see processing time for all affected services has returned to normal as of 20 minutes ago. May 26, 10:08 UTC Monitoring - We've identified the problem and have completed the remediation steps, we are now monitoring as service resumes. May 26, 09:56 UTC Identified - We're observing increased laten…
May 20, 17:39 UTC Resolved - The incident is resolved May 20, 17:26 UTC Monitoring - We are seeing recovery across affected customers and continue to monitor May 20, 17:06 UTC Identified - We have identified the issue and applied mitigations and are monitoring recoveryWe have determined that only a subset of customers are affected by the notification latency. May 20, 16:40 UTC Investigating - We a…
May 15, 07:35 UTC Resolved - Processing of the backlog is complete. May 15, 06:51 UTC Monitoring - Ingestion of Test Engine execution data from an internal queue to a data store stalled, has been resumed, and is working through the backlog. Visibility of test executions from the past hour hours will be delayed for approximately a further one hour.This has been a recurring issue; an architectural c…
May 13, 15:34 UTC Resolved - Additional capacity was added to our redis caches. This triggered a failover between UTC 15:10 - 15:14 and there was a spike of errors on the REST and GraphQL APIs. Customers would have seen some errors in the Buildkite UI during this period as well. We have been monitoring the situation since then and things have returned to baseline. May 13, 15:14 UTC Investigating -…
May 12, 16:09 UTC Resolved - The fix was successful and the backlog has now been cleared. May 12, 12:59 UTC Monitoring - We are currently experiencing delayed processing of Test Engine data. We have identified and applied a fix for the issue but are expecting to continue to experience delays while we clear the ingestion backlog
May 9, 04:34 UTC Resolved - The upstream AWS incident in us-east-1 has been resolved by AWS, and all Buildkite services are operating normally. No further customer impact is expected. We appreciate your patience during this incident. May 8, 08:07 UTC Monitoring - Despite the ongoing AWS incident, our own services are now stable. We are continuing to monitor our services closely, and are ready for…
May 8, 23:42 UTC Resolved - The delayed backlog is now cleared and Test Engine ingestion is operating normally. May 8, 21:10 UTC Investigating - We are currently experiencing delayed processing of Test Engine data. We have identified and applied a fix for the issue but are expecting to continue to experience delays while we clear the ingestion backlog. At the current processing rate we expect the…
May 8, 13:06 UTC Resolved - The Test Engine backlog is now cleared and operating normally. May 8, 12:20 UTC Monitoring - We are currently experiencing delayed processing of Test Engine data. We have identified and applied a fix for the issue but are expecting to continue to experience continued processing delays while we clear the ingestion backlog. At the current processing rate we expect the bac…
May 8, 00:13 UTC Resolved - All customer workloads have now recovered. May 7, 23:50 UTC Monitoring - We have now monitoring the incident. We are seeing most customers have recovered, and some showing signs of recovery. May 7, 23:13 UTC Identified - We've identified the issue and are working on applying mitigations. At this time we can confirm inbound and outbound webhooks, and notifications are de…
May 7, 09:37 UTC Resolved - We've reverted a change that caused stale environment variables provided to acquire job used in Hosted Agents, agent-stack-k8s and other agent implementations using acquire job. May 7, 09:24 UTC Identified - We're currently seeing recovery at 50% rate. We'll provide next update soon. May 7, 09:09 UTC Update - We've identified issue with job acquiring endpoint. We're rol…
May 6, 05:26 UTC Resolved - Processing of test execution ingestion data has successfully caught up. May 6, 04:21 UTC Monitoring - We've identified the issue and the system is currently processing the backlog of test executions May 6, 03:57 UTC Investigating - A process writing test results to our Test Engine data store stalled, we've restarted the process and are seeing it catching up. We expect t…
May 4, 19:36 UTC Resolved - Notifications continue to be delivered without delay for the previously affected subset of customers. This incident is resolved. May 4, 18:54 UTC Monitoring - Applied remediations have resolved the previous notification delays affecting a subset of our customers. We're continuing to monitor the affected services for stability. May 4, 18:06 UTC Identified - We've identif…
May 4, 06:30 UTC Resolved - An increase in requests has lead to the API service being temporarily saturated. We have updated rate limits to ensure this doesn't re-occur and will add further resources if necessary May 4, 06:02 UTC Investigating - We're observing increased latency and error rates in the Agent API for a subset of our customers. We're currently investigating and will provide status up…
May 1, 03:30 UTC Resolved - This incident has been resolved. May 1, 02:52 UTC Monitoring - We've corrected the issue that caused this disruption and normal service has been restored. We are monitoring the situation now. May 1, 02:41 UTC Identified - We've identified a service change that is causing a service disruption. We are reverting this change.
Apr 29, 18:48 UTC Resolved - We have confirmed that latency and error rates have returned to normal for impacted customers. Apr 29, 18:18 UTC Monitoring - We have identified and fixed the issue with the underlying database for a subset of customers. We are now monitoring the issue. Apr 29, 17:43 UTC Investigating - We're observing increased latency and error rates for a subset of our customers. We…
Apr 28, 19:16 UTC Resolved - Previously elevated loads with Hosted Agents dispatch have fully recovered. Apr 28, 18:45 UTC Monitoring - We have mitigated the issue causing increased Hosted Agents dispatch latency and intermittent timeout errors for a subset of customers. We identified abnormal workload activity that was placing elevated load on a supporting service, and have now blocked that activ…
Apr 22, 22:59 UTC Resolved - The issue is resolved. Apr 22, 22:44 UTC Monitoring - We have rolled back a change on the remote MCP server that was contributing to authentication failures. Apr 22, 22:07 UTC Update - We are continuing to investigate errors when authenticating to the remote MCP server. Apr 22, 21:19 UTC Investigating - We are currently investigating reports of authentication failures…
Apr 22, 05:07 UTC Resolved - The backlog has been cleared and all systems are fully operational. Thank you for your patience. Apr 22, 02:32 UTC Monitoring - We noticed a lag in data processing, but our systems are operational and currently working through the backlog. We expect to be fully caught up within the next couple of hours.
Apr 8, 23:12 UTC Resolved - We experienced an issue that caused a brief increase in errors for the Agent API. This also impacted latency for notifications. All notifications were stored in a queue and processed. Latency is now back to normal. Apr 8, 22:40 UTC Monitoring - We have identified and fixed the issue. We are monitoring and seeing signs of improvement. Apr 8, 22:26 UTC Investigating - We'…
Mar 31, 08:34 UTC Resolved - We have deployed the fix and we have confirmed customer builds are working. If you encounter any further issues please contact support. Mar 31, 08:15 UTC Identified - We have identified the issue and are rolling out a fix. Mar 31, 07:51 UTC Investigating - We have received reports from customers that they are unable to start builds on Hosted Agents. Their builds are im…