1
Detection
0-15 min
2
Triage
15-45 min
3
Response
1-4 hrs
4
Resolution
4-24 hrs
5
Post-Mortem
48 hrs

1. Detection & Intake

Automated + Manual

Incidents are surfaced through automated monitoring (CMS logs, CDN errors, security scanners) or manual reports from staff, editors, or external readers. All reports are logged with timestamps and initial metadata.

Primary Triggers

  • CMS deployment failure or rollback
  • CDN latency spikes (>500ms)
  • Fact-check flag from editorial board
  • DDoS or unauthorized access attempt
  • Reader/whistleblower submission

Responsible Roles

  • On-Call Editor / Duty Desk
  • Platform Engineering
  • Security Operations Center

2. Triage & Classification

Severity Scoring

The duty desk classifies incidents by type (Editorial, Infrastructure, Security, Compliance) and assigns a severity level (Critical, High, Medium, Low). SLA timers activate based on classification.

Severity Matrix

  • Critical: Live broadcast down, major factual error on front page
  • High: CMS outage, data breach risk, multi-article misprint
  • Medium: Single page render issue, minor attribution error
  • Low: Typographical error, non-breaking UI glitch

Escalation Paths

  • Editorial → Editor-in-Chief / Legal
  • Technical → CTO / DevOps Lead
  • Security → CISO / IR Team

3. Response & Containment

Active Mitigation

Specialized teams execute playbooks to contain impact. Editorial issues trigger immediate pulls/corrections. Technical incidents may require feature flags, rollbacks, or traffic rerouting.

Editorial Playbook

  • Issue immediate correction notice
  • Pause related syndication feeds
  • Brief social & PR teams

Technical/Security Playbook

  • Isolate affected services
  • Deploy hotfix or rollback
  • Enable maintenance mode if critical

4. Resolution & Verification

Validation Phase

Once mitigated, the incident undergoes verification by independent reviewers. Corrections are published, systems are restored to full capacity, and stakeholders receive status updates.

Verification Checklist

  • Cross-departmental sign-off
  • Load testing restored baselines
  • Fact-check audit passed

Communication

  • Internal memo to newsroom
  • Public transparency log update
  • Reader notification (if applicable)

5. Post-Incident Review

Continuous Improvement

A blameless post-mortem is conducted within 48 hours. Root cause analysis, timeline reconstruction, and actionable improvements are documented and tracked in our knowledge base.

Deliverables

  • Root Cause Analysis (RCA) document
  • Timeline & impact assessment
  • Preventive action items with owners

Tracking

  • Action items logged in Jira/Linear
  • Quarterly workflow audit
  • Metrics: MTTR, MTTD, recurrence rate

Recent Incidents

ID Type Severity Status Assigned Timestamp
#INC-4829 Editorial High ● In Progress Editorial Desk 2025-11-14 09:12 UTC
#INC-4828 Infrastructure Critical ● Resolved Platform Eng 2025-11-13 14:45 UTC
#INC-4827 Security Medium ● Closed IR Team 2025-11-12 11:30 UTC
#INC-4826 Editorial Low ● Closed Copy Desk 2025-11-11 16:20 UTC

Submit New Incident