WikiOasis

Report a problemSubscribe to updates
Powered by
Privacy policy

·

Terms of service
Write-up
Search degraded for some wikis
Partial outage
View the incident
Summary

Search was degraded due to the deployment of CirrusSearch globally. As a whole, the impact was minimal, affecting search across most wikis, with some having search entirely unavailable, and others having degraded performance due to the cluster being under heavy load.

Impact
  • Users on some wikis, ~90% at the start which decreased quickly as the incident went on, would have been unable to search successfully.

  • Users on the remaining wikis which did have search available may have experienced a slight degradation in the speed of search, due to the search cluster being under heavy load from the ongoing indexing job.

Incident Timeline

Time

Event

2026-08-21

09:30:00

Impact started at ​ Custom timestamp "Impact started at" occurred

10:46:24

Incident reported by Thomas ​ Thomas reported the incident Severity: SEV2 Status: Investigating

10:55:32

Identified at ​ Custom timestamp "Identified at" occurred

10:55:32

Status changed from Investigating → Fixing ​ Thomas shared an update Status: Investigating → Fixing Cause has been identified to be ongoing heavy indexing job — not a lot can be done beyond sizing up, but it would appear the cluster is having slightly degraded performance — indexing is probably nearly finished

11:06:56

Status changed from Fixing → Monitoring ​ Thomas shared an update Status: Fixing → Monitoring By and large, not a lot can be done — the index job will complete, it appears to be not having as much of an impact on latency or search performance now

11:06:56

Fixed at ​ Custom timestamp "Fixed at" occurred

11:06:56

Impact ended at ​ Custom timestamp "Impact ended at" occurred

11:14:09

Update shared ​ Thomas shared an update Maximum number of shards was hit, increased to 6000 per active data node. Cluster size may need to grow

11:26:40

Incident resolved and entered the post-incident flow ​ Thomas shared an update Status: Monitoring → Documenting Opensearch has returned to norma

3h later

14:51:08

Documented at ​ Custom timestamp "Documented at" occurred

14:51:08

Status changed from Documenting → Reviewing ​ incident.io (incident lifecycle) shared an update Status: Documenting → Reviewing

14:51:18

Reviewed at ​ Custom timestamp "Reviewed at" occurred

14:51:18

Status changed from Reviewing → Finalising ​ incident.io (incident lifecycle) shared an update Status: Reviewing → Finalising

14:52:05

Incident closed ​ incident.io (incident lifecycle) shared an update Status: Finalising → Closed

Contributors
  • A background job to index search across all wikis as part of a deployment for CirrusSearch globally caused this incident

  • The incident effects were compounded by the OpenSearch cluster running out of available shards, meaning that some wikis were unable to successfully index and indexing jobs had to be retried.

Root Cause

The deployment of CirrusSearch was not done in a controlled manner, meaning that all wikis had to index at the same time, instead of a rolling deployment and indexing which would've taken down search on fewer wikis at a time, rather than all being queued at once.

Mitigators

Having clear observability setup for the OpenSearch cluster meant that we were easily able to check the status of indexing, and check that the cluster was not overloaded. This allowed us to see at times when the cluster went into a yellow state.

Learnings and risks

At times, the OpenSearch cluster may be too small to be able to handle large indexing jobs without service degradation.