DEV Community

Site Reliability Engineering

Site Reliability Engineering principles, practices, and culture.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Netdata Alerts in Week One: Which to Keep, Tune or Silence, and How to Route Them to Slack, Discord and ntfy

Netdata Alerts in Week One: Which to Keep, Tune or Silence, and How to Route Them to Slack, Discord and ntfy

Comments 1
12 min read
Top Kubernetes Production Incident Scenarios & Diagnostic Runbooks (2026 Edition)

Top Kubernetes Production Incident Scenarios & Diagnostic Runbooks (2026 Edition)

1
Comments
5 min read
Teaching SRE in Resistant Organisations: A Phased Influence Playbook for Practitioners

Teaching SRE in Resistant Organisations: A Phased Influence Playbook for Practitioners

Comments
12 min read
What should trigger an autonomous agent in production?

What should trigger an autonomous agent in production?

3
Comments 1
7 min read
Ringbolt: an on-call line that has to hear you say it

Ringbolt: an on-call line that has to hear you say it

Comments
3 min read
Beyond Uptime Checks: Meaningful Infrastructure Monitoring

Beyond Uptime Checks: Meaningful Infrastructure Monitoring

Comments
2 min read
Why I Built Compute Central: A Practical DevOps, SRE, and Cloud Knowledge Base

Why I Built Compute Central: A Practical DevOps, SRE, and Cloud Knowledge Base

Comments
1 min read
One instance was bad at its job and nothing was designed to notice

One instance was bad at its job and nothing was designed to notice

Comments
2 min read
Latest N Is Not a Time Window

Latest N Is Not a Time Window

Comments 1
4 min read
The cold start we had never practised

The cold start we had never practised

Comments
2 min read
Kubernetes Troubleshooting: What to Check Before You Restart a Pod

Kubernetes Troubleshooting: What to Check Before You Restart a Pod

Comments
2 min read
Nothing could start without the dependency we had filed as optional

Nothing could start without the dependency we had filed as optional

Comments 1
2 min read
Backyard Endurance OS: Designing Zero-Loss Telemetry Ingestion for Athletes and Distributed Systems

Backyard Endurance OS: Designing Zero-Loss Telemetry Ingestion for Athletes and Distributed Systems

Comments 1
5 min read
Node.js Consumer Capacity: Webhook Delivery, Rate Limits, and Dead-Letter Decisions

Node.js Consumer Capacity: Webhook Delivery, Rate Limits, and Dead-Letter Decisions

1
Comments
6 min read
Top Kubernetes Production Incident Scenarios & Diagnostic Runbooks (2026 Edition)

Top Kubernetes Production Incident Scenarios & Diagnostic Runbooks (2026 Edition)

1
Comments
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.