-
Assume the Incident Commander role for all P1 and P2 incidents: own the bridge call, drive resolution, and coordinate cross-functional resolver groups.
-
Establish roles on incident bridges (scribe, technical lead, communications lead) and enforce time-boxed troubleshooting with 30-minute checkpoints to avoid stagnation.
-
Make escalation decisions: page on-call engineers, engage leadership or vendors, notify compliance, or execute rollback/failover/service-degradation procedures.
-
Send initial stakeholder notifications within the SLA and maintain a consistent cadence: technical details for engineering; business impact for leadership.
-
Coordinate with Compliance and Regulatory teams for impact notifications and initiate customer-facing communications (app banners, status page updates) when required.
-
Enforce monitoring coverage requirements for all Caesars Digital production services, ensure no service goes to production without adequate monitoring and alerting in place.
-
Continuously tune alert thresholds based on feedback from Digital System Support Engineers, reducing false positives and alert fatigue while driving the alert signal-to-noise ratio above 80% actionable.
-
Implement alert deduplication, correlation, and suppression rules to ensure Engineers receive clean, actionable signals rather than noise that degrades response effectiveness.
-
Define and maintain alert severity standards that clearly distinguish P1 vs. P2 vs. P3 vs. informational alerts, ensuring consistent classification across all services.
-
Review "missed detection" findings from the Major Incident Manager's Post-Incident Reviews and build new monitoring coverage to prevent recurrence of undetected issues.
-
Ensure every alert in the ecosystem links to a documented runbook with clear response procedures that Engineers can execute independently.
-
Build and own end-to-end customer journey monitoring covering the critical user flows: Registration Deposit Bet Placement Bet Settlement Withdrawal, with defined thresholds for success rates and drop-off alerts.
-
Design and implement real-time revenue monitoring dashboards tracking deposit/withdrawal volumes, payment gateway health, and transaction success rates with anomaly detection against expected baselines.
-
Build revenue impact calculation models for use during major incidents, enabling the team to quantify business impact in dollar terms.
-
Define and maintain business KPI dashboards monitoring operational metrics including handle, active users, concurrent sessions, bet volume per minute, and Caesars Rewards pipeline health.
-
Create executive-visible business health dashboards that provide real-time situational awareness during high-revenue events and peak traffic periods.
-
Conduct monthly monitoring audits to assess coverage completeness, alert quality, and identify stale or orphaned alerts for decommissioned services.
-
Build synthetic monitoring for critical customer journeys to validate service availability and performance from the customer's perspective.
-
Collaborate with the Problem Manager to support event readiness by building enhanced dashboards, lowering detection thresholds, and adding event-specific synthetic monitors per the readiness plan.
-
Report on noisy alert sources monthly and drive engineering teams to fix the root causes generating non-actionable alerts.
-
Partner with product and engineering teams to agree on monitoring thresholds and ensure observability is built into the development lifecycle.