Red - System Down

Incident Report for PEAK 15 Systems

Postmortem

Service Incident Postmortem

Incident date: August 6, 2026
Incident window: Approximately 1:05–1:20 AM PDT

Summary

At approximately 1:05 AM PDT, the production database server was inadvertently shut down while maintenance was being performed on a separate server.

Because our services depend on the database platform, customers experienced a service interruption or critically degraded functionality while the database server was offline and SQL Server was unavailable.

Our team investigated the interruption, identified that the database server had been shut down, and initiated recovery. The server completed its startup process at approximately 1:20 AM PDT, after which database connectivity and dependent services were restored.

Timeline

  • Approximately 1:05 AM PDT: The production database server was inadvertently shut down.
  • 1:05–1:20 AM PDT: Our team investigated the service interruption, identified the cause, and initiated recovery of the database server.
  • Approximately 1:20 AM PDT: The database server completed startup, SQL Server became available, and dependent application services recovered.

Root Cause

The incident was caused by human error during maintenance activity involving another server. An action intended for the maintenance target was mistakenly applied to the production database server.

Corrective Actions

We are reviewing our maintenance procedures and administrative safeguards to reduce the likelihood of a similar error. Corrective actions will include strengthening server-identification and verification steps before disruptive actions are performed, along with reviewing available Azure controls that can provide additional protection for critical production resources.

We apologize for the disruption and appreciate your patience while service was restored.

Posted Aug 06, 2026 - 11:02 PDT

Resolved

Our database server was inadvertently taken offline during maintenance being performed on a separate server. This resulted in a temporary service interruption.

Service has now been fully restored. We are conducting a formal review of the incident and will complete a postmortem tomorrow morning to determine the contributing factors and identify corrective actions to prevent a recurrence.

We apologize for the disruption and appreciate your patience.
Posted Aug 06, 2026 - 01:36 PDT

Identified

The issue has been identified and a fix is being implemented.
Posted Aug 06, 2026 - 01:20 PDT

Investigating

We are currently experiencing a system-wide outage affecting all customers and services.

Our engineering team is actively investigating the issue and working to restore service as quickly as possible. We understand the urgency of this situation and are treating it with the highest priority.

At this time, we do not have an estimated time to resolution. We will provide updates here as soon as more information becomes available.

Thank you for your patience while we work to resolve this issue.
Posted Aug 06, 2026 - 01:08 PDT