Search was degraded due to the deployment of CirrusSearch globally. As a whole, the impact was minimal, affecting search across most wikis, with some having search entirely unavailable, and others having degraded performance due to the cluster being under heavy load.
Users on some wikis, ~90% at the start which decreased quickly as the incident went on, would have been unable to search successfully.
Users on the remaining wikis which did have search available may have experienced a slight degradation in the speed of search, due to the search cluster being under heavy load from the ongoing indexing job.
Time | Event |
|---|---|
2026-08-21 | |
09:30:00 | Impact started at Custom timestamp "Impact started at" occurred |
10:46:24 | Incident reported by Thomas Thomas reported the incident Severity: SEV2 Status: Investigating |
10:55:32 | Identified at Custom timestamp "Identified at" occurred |
10:55:32 | Status changed from Investigating → Fixing Thomas shared an update Status: Investigating → Fixing Cause has been identified to be ongoing heavy indexing job — not a lot can be done beyond sizing up, but it would appear the cluster is having slightly degraded performance — indexing is probably nearly finished |
11:06:56 | Status changed from Fixing → Monitoring Thomas shared an update Status: Fixing → Monitoring By and large, not a lot can be done — the index job will complete, it appears to be not having as much of an impact on latency or search performance now |
11:06:56 | Fixed at Custom timestamp "Fixed at" occurred |
11:06:56 | Impact ended at Custom timestamp "Impact ended at" occurred |
11:14:09 | Update shared Thomas shared an update Maximum number of shards was hit, increased to 6000 per active data node. Cluster size may need to grow |
11:26:40 | Incident resolved and entered the post-incident flow Thomas shared an update Status: Monitoring → Documenting Opensearch has returned to norma |
3h later | |
14:51:08 | Documented at Custom timestamp "Documented at" occurred |
14:51:08 | Status changed from Documenting → Reviewing incident.io (incident lifecycle) shared an update Status: Documenting → Reviewing |
14:51:18 | Reviewed at Custom timestamp "Reviewed at" occurred |
14:51:18 | Status changed from Reviewing → Finalising incident.io (incident lifecycle) shared an update Status: Reviewing → Finalising |
14:52:05 | Incident closed incident.io (incident lifecycle) shared an update Status: Finalising → Closed |
A background job to index search across all wikis as part of a deployment for CirrusSearch globally caused this incident
The incident effects were compounded by the OpenSearch cluster running out of available shards, meaning that some wikis were unable to successfully index and indexing jobs had to be retried.
The deployment of CirrusSearch was not done in a controlled manner, meaning that all wikis had to index at the same time, instead of a rolling deployment and indexing which would've taken down search on fewer wikis at a time, rather than all being queued at once.
Having clear observability setup for the OpenSearch cluster meant that we were easily able to check the status of indexing, and check that the cluster was not overloaded. This allowed us to see at times when the cluster went into a yellow state.
At times, the OpenSearch cluster may be too small to be able to handle large indexing jobs without service degradation.