Back

Incident history

August
Dashboard Analytics Behind
Resolved

A postmortem for the incident is below. Please reach out to Support for further questions.

 

Overview

On August 18th between 2:00pm ET and 8:00pm ET Dashboard analytics experienced a delay of up to 5 hours. Customers viewing analytics during this window may have seen incomplete data for the current day. No analytics data was lost or corrupted during this incident.

 

Incident Summary

Admiral’s analytics ingestion process handles messages from our JavaScript client asynchronously, normally processing them within seconds. On August 18th, an unannounced update to an underlying third-party infrastructure component caused this process to fall behind processing events. Once we identified the change, a custom workaround was deployed, and the system began to automatically process the backlogged analytics.

 

Root Cause Analysis

Admiral relies on enterprise service providers to support our platform’s infrastructure. On the day of the incident, an underlying infrastructure provider rolled out a change to how internal identifiers are formatted. This format change unexpectedly triggered internal errors during the processing of events. These errors led to a bottleneck between our platform and the provider. The reduced delivery was a downstream effect and not the underlying cause which took several hours for our team to uncover.

 

Mitigation Steps

We designed our analytics ingestion to prioritize accuracy over performance and the undelivered events were not lost and instead were being queued until they could be successfully processed. Because the format change occurred deep within the provider, it took our engineering team time to trace and isolate the issue. Once identified, our team deployed a custom configuration to bypass the provider’s errors and successfully restore the normal flow of events. At that time, analytics processing was several hours behind and took approximately 90 minutes to safely process the queued data and catch up to real-time.

 

Preventive Measures

We worked with the provider to deploy a permanent update and prevent a recurrence. Additionally, we have rolled out improved metrics to monitor provider queue health. Finally, we are working closely with our infrastructure provider to improve notifications and testing surrounding their future system updates.

August 18 at 03:43 PM EDT

Resolved after 3h 50m

July

No incidents reported

June

No incidents reported