What Is Site Reliability Engineering? A Solo Dev's Take
Summary
What is site reliability engineering? It applies software engineering to operations: you pick a measurable reliability target (an SLO), track it with an indicator (an SLI), and use the error budget, which is 100% minus the target, to decide whether to ship features or fix reliability. Large teams add on-call rotations and postmortems. A solo builder needs far less: one indicator, one loose target, a health check, and a short postmortem file.
What is site reliability engineering? It is the practice of treating operations as a software problem: you set a measurable reliability target, track how often you miss it, and let that number decide whether you ship features or fix things. Google coined the term, but the idea works at any size. For a solo dev with one app and a handful of users, it means deciding, in advance, how much downtime you will accept.
Here is what gets stuck in practice: most side projects have no such number. They are "up" until someone emails you. This guide explains SRE in plain terms, then shrinks it to something one person can run on a weekend.
So what is site reliability engineering, really?
The short version comes from Google's SRE book: SRE is what you get when you ask a software engineer to design an operations team. Instead of people manually restarting servers, you write code that does it, and you measure the results.
Three ideas carry most of the weight. Reliability is a feature with a target, not a vibe. Repetitive manual work (the book calls it toil) is a bug to be automated away. And failure is expected, so you plan how to learn from it instead of how to avoid blame.
DevOps and SRE overlap a lot. The practical difference: DevOps is a culture of shipping and running code together, while SRE gives you specific tools to argue about it with numbers. Nobody needs to pick a camp to use the numbers.
SLI, SLO, SLA: three acronyms, one idea each
An SLI (service level indicator) is something you measure. For a web app, that is usually the share of requests that succeed, or the share that answer in under 300 ms. Pick one or two. Not twelve.
An SLO (service level objective) is the target you set on that measurement, for example "99.9% of requests succeed over 30 days". This is internal. It is a promise to yourself.
An SLA (service level agreement) is a contract with a customer, usually with refunds attached if you miss it. If nobody has paid you for uptime, you do not have an SLA yet, and you should not write one by accident in your pricing page copy.

Error budgets: the part that changes how you work
The error budget is simply 100% minus your SLO. A 99.9% target over 30 days leaves about 43 minutes of allowed failure. That is the budget you can spend on risky deploys, migrations, and experiments.
The useful bit is the rule attached to it. While budget remains, you ship. When it is gone, you stop shipping features and fix reliability until it recovers. The book's chapter on risk frames it as a way to settle the argument between "move faster" and "don't break things" with a shared number instead of a negotiation.
It also makes a blunt point worth repeating: 100% is almost never the right target. Users on flaky mobile networks cannot tell 99.99% from 99.9%, and each extra nine costs far more than the last. It is not a perfect target. It is a workable one.
If you want the concrete workflow for picking indicators and writing the first objective, the SRE workbook on implementing SLOs is the most practical chapter and runs short.
What does an SRE actually do all day?
At a large company, an SRE splits time between on-call, incident response, capacity planning, and automation work. Google's book says to cap operational load at around half of an engineer's time so the rest goes to engineering. That cap is a design constraint, not a nice-to-have.
The recurring work looks like this:
Defining SLOs with product teams and alerting on budget burn, not on every blip
Running on-call rotations and writing runbooks so the 3 a.m. page has a checklist
Leading incident response, then writing a blameless postmortem
Automating anything done manually more than a few times
Reviewing launches for reliability risks before they hit production
Feature flags are a good example of the mindset. A flag lets you turn off a bad release in seconds without a redeploy, which protects the budget. Whether you need a hosted tool for that depends on how many flags you have: an environment variable can do the job for the first three.
Do you need SRE if you are building alone?
Three reasons to do it, one reason not to.
The reasons to do it: you will get paged by your own project eventually; a written target stops you from over-engineering; and "I run production with an SLO" reads well when someone is deciding whether to hire or trust you. The reason not to: if your project has no users yet, you are practising for a problem you do not have. Build first, measure when someone depends on it.
So skip the full apparatus. Do not hire a pager service for a hobby app, do not run Kubernetes to feel serious, and do not write a 10-page incident template. Worth doing if you have even ten real users: a health check, one SLO, and a place to look when it breaks.

A one-evening SRE setup for a side project
Here is the minimum that tends to hold up at 10 000 users, too. It takes an evening. Tested. Not optimal. Here is why we do it anyway: it is boring, and boring survives.
First, pick one SLI: the share of HTTP requests returning something other than a 5xx. Second, set the SLO at 99.5% over 30 days, which is about 3.6 hours of budget. Start loose, tighten later. Third, add an external uptime check hitting a real health endpoint.
A health endpoint worth having checks the things that actually break, not just that the process is alive:
// GET /healthz
app.get("/healthz", async (req, res) => {
try {
await db.query("select 1"); // database reachable
await cache.ping(); // cache reachable
res.status(200).json({ ok: true });
} catch (err) {
res.status(503).json({ ok: false });
}
});Fourth, send alerts to somewhere you will see them, a phone push rather than an inbox. Fifth, keep a plain text file called postmortems.md and add four lines after every outage: what happened, why, how long, what changes. That file is the most underrated reliability tool you will own.
Where AI tools help and where they do not
Assistants are decent at the toil side: writing the health check, the Terraform for an uptime monitor, a first draft of a runbook, or a script that parses logs for the 5xx rate. They are bad at deciding your SLO, because that is a product call about what your users tolerate.
During an incident, be careful. An assistant can suggest a plausible fix that you then apply to production at 2 a.m. without reading it. Use it to explain a stack trace or draft the postmortem afterward, and keep the rollback decision with a human who is awake.
Code quality gates belong in this picture too. Many outages trace back to a change that looked harmless in review. Static analysis in CI will not give you reliability, but it removes a class of dumb mistakes before they cost budget.
How do you write a postmortem that someone will read?
Keep it short and blameless. Blameless does not mean nobody is accountable. It means you ask what in the system allowed the mistake, not who made it. When you are the only person on the team, this matters more than it sounds, because the temptation is to feel bad and skip the writeup.
Four headings are enough: what happened, why it happened, how long users were affected, and what changes. The last one should name an action you will really do, with a date. "Be more careful" is not an action. "Add a staging check for migrations before the next release" is.
Over a year, that file turns into a map of your weak spots. The rule after three outages caused by the same expired certificate: if the same cause shows up twice, automate it. It took one evening to add renewal monitoring, and it has not recurred.
Alert on burn rate, not on every blip
A common beginner mistake is paging yourself whenever a single request fails. You will mute the channel within a week. A better signal is burn rate: how fast you are spending the error budget compared with the pace that would exhaust it exactly at the end of the window.
For a side project, keep it crude. Send a push notification if the error rate over the last 10 minutes is above 5%, and an email summary if the weekly rate is trending toward missing the SLO. That gives you two levels: wake up now, or look on Monday. Anything fancier can wait until users complain about the first level.
Is SRE a good career path for a dev who likes building?
It depends on what you enjoy. SRE roles suit people who like systems, failure modes, and automation more than shipping UI. If you get satisfaction from making the pager quieter, you will like it. If you want to see users click your new feature, you probably will not.
The path in is usually through backend or infrastructure work: you start owning a service's deploys and alerts, then take on more of its reliability. Skills that carry over: Linux basics, networking, one cloud provider, a monitoring stack, and the habit of writing things down. Side projects are a fair place to build all of them, because you are the whole team.

What to build next
Pick one thing you already run, even a free-tier app, and write down its SLI, SLO, and the exact action you will take when the budget runs out. If you cannot say what you would stop doing, the number is decoration. Which of your projects would you notice was down before a user told you?