A postmortem for the incident is below. Please reach out to Support for further questions.
Overview
On August 18th between 2:00pm ET and 8:00pm ET Dashboard analytics experienced a delay of up to 5 hours. Customers viewing analytics during this window may have seen incomplete data for the current day. No analytics data was lost or corrupted during this incident.
Incident Summary
Admiral’s analytics ingestion process handles messages from our JavaScript client asynchronously, normally processing them within seconds. On August 18th, an unannounced update to an underlying third-party infrastructure component caused this process to fall behind processing events. Once we identified the change, a custom workaround was deployed, and the system began to automatically process the backlogged analytics.
Root Cause Analysis
Admiral relies on enterprise service providers to support our platform’s infrastructure. On the day of the incident, an underlying infrastructure provider rolled out a change to how internal identifiers are formatted. This format change unexpectedly triggered internal errors during the processing of events. These errors led to a bottleneck between our platform and the provider. The reduced delivery was a downstream effect and not the underlying cause which took several hours for our team to uncover.
Mitigation Steps
We designed our analytics ingestion to prioritize accuracy over performance and the undelivered events were not lost and instead were being queued until they could be successfully processed. Because the format change occurred deep within the provider, it took our engineering team time to trace and isolate the issue. Once identified, our team deployed a custom configuration to bypass the provider’s errors and successfully restore the normal flow of events. At that time, analytics processing was several hours behind and took approximately 90 minutes to safely process the queued data and catch up to real-time.
Preventive Measures
We worked with the provider to deploy a permanent update and prevent a recurrence. Additionally, we have rolled out improved metrics to monitor provider queue health. Finally, we are working closely with our infrastructure provider to improve notifications and testing surrounding their future system updates.
Resolved after 3h 50m
No incidents reported
No incidents reported