PayCrunch Research · The exact AI playbook for your profession, sourced to the U.S. Bureau of Labor Statistics

PayCrunch AI Playbook · Technology

How a site reliability engineer argues for more ground

$272,670top of the range in California · middle $135,980 / yr
AI augments this role

Site Reliability Engineers in the United States earn a median of $135,980 a year. Pay starts near $82,460. Pay reaches $272,670 at the top of the range in California, the best-paying state for this work among those with at least 500 people in the job.

Source: U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (Software Developers, SOC 15-1252). Last checked 9 September 2026.

Entry level
$82,460
Top of the range · California
$272,670
Education
Bachelor's degree in Computer Science
Lower disruption Higher exposure AI augments this role
Entry · $82,460 Top of range · $272,670 (California) Middle $135,980

Wages — U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (Software Developers). Top of the range is the highest state-level figure among states with at least 500 people in the job. AI-impact rating is PayCrunch's editorial assessment. Updated September 2026.

🆕 New & Trending AI Tools for Site Reliability EngineerReviewed September 2026

We track new AI-tool launches every week and refresh this list — here’s what’s gaining traction for Site Reliability Engineer work right now.

Claude CodeNEWFree / usage-based

Terminal coding agent that reads your repo, runs tests, and ships multi-file changes.

How a Site Reliability Engineer uses it: describe a feature and let it implement and test it across the codebase

OpenAI CodexNEWIncl. w/ ChatGPT plans

Agent that runs longer, deterministic multi-step coding jobs on its own.

How a Site Reliability Engineer uses it: delegate a well-defined build or migration and review the finished result

WindsurfNEWFree / $15 mo

Agentic IDE that keeps context across a whole project.

How a Site Reliability Engineer uses it: make large, coordinated changes without losing track of the codebase

AWS KiroNEWPreview / see site

Spec-driven coding agent that turns written specs into working code.

How a Site Reliability Engineer uses it: write the spec first and let it build to that spec

NotebookLMNEWFree / $7.99 mo

Google tool that answers questions grounded only in the documents you give it — with citations.

How a Site Reliability Engineer uses it: load your own manuals, policies, or PDFs and ask questions that stay accurate to the source

CursorFree / $20 mo

AI-native code editor that edits across an entire project.

How a Site Reliability Engineer uses it: describe a change in plain English and let it rewrite and refactor whole files

GitHub Copilot (Agent Mode)$10–19 mo

AI pair-programmer built into VS Code and GitHub that now completes multi-step tasks.

How a Site Reliability Engineer uses it: hand off a task and have it plan, edit multiple files, and open a pull request

ChatGPTFree / $20 mo

The most-used AI assistant — writing, analysis, research, and images from a plain-language chat.

How a Site Reliability Engineer uses it: draft emails and documents, summarize long files, and get instant answers to on-the-job questions

ClaudeFree / $20 mo

AI assistant known for careful writing, long-document analysis, and coding.

How a Site Reliability Engineer uses it: analyze big reports or spreadsheets and turn messy notes into clean, finished writing

The error budget that is almost spent

The checkout service has almost no error budget left for the month, and a deploy from the afternoon is the likely cause. A site reliability engineer opens the dashboard before opening the chat: error rate, latency, saturation on the database, and the release marker sitting on the timeline. The feature team wants to keep shipping. The budget says the service has already spent the unreliability the organization agreed to accept. The engineer pages the owner of the change, starts a rollback checklist, and writes the first note of what customers would have felt. The job in that hour is the reliability of a running service.

Site reliability engineers live on the production path. They define what "up" means with the people who own the product, they watch the signals that prove it, and they intervene when the service drifts. They sit with feature teams during design long enough to ask how a failure will be detected, and they sit with support when a customer hits the symptom first. The tools are dashboards, alerts, deploy systems, tracing, and a log search that has to work under stress. The decision is often whether to roll back, to shed load, or to hold the next release until the budget recovers. A pipeline that builds artifacts matters because it is how change arrives. The engineer's name is on the service after the artifact is live.

The week also holds quieter work. You delete an alert that pages for a condition nobody can act on. You turn a manual restart into a safe control with a permission and a log. You read a design for a new dependency and ask what the service does when that dependency times out. You write the incident timeline while memories are still fresh. Places include a company office, a home desk with the laptop that receives pages, and the virtual room where an incident is run. The people are feature owners, an incident lead, support, and sometimes a customer-facing manager who needs a sentence that is true.

Incidents, toil, and the path a release takes

An error budget is an agreement. The service has a target for reliability over a window, and the budget is the room left inside that target. When the budget is healthy, feature work proceeds. When it is thin, reliability work takes the front of the queue: a bad deploy path, a dependency with no fallback, a capacity cliff. You help the team pick the signal, the window, and the consequence, and you keep the budget visible so the argument happens before the outage. You do not need a private dialect to do this. You need a number the product owner accepts and a page that fires only when a person should act.

Incidents are the public half of the job. You join or lead, you protect the timeline from guesswork, you assign the next check, and you decide when the service is stable enough to stop the bridge. Afterward you write what happened, what customers experienced, and which change will keep this class of failure from repeating. Blame aimed at a single commit teaches people to hide. A repair aimed at the path teaches the next deploy. Toil is the other half: repeated manual work that does not need human judgment. Restarting a process by hand, copying a token between tools, clicking through a failover you have memorized. You count that work by how often it steals an afternoon, and you replace the worst of it with automation that has a clear owner.

The production path is the sequence a change actually walks: review, build, progressive rollout, observation, and a fast way back. A site reliability engineer cares that each step leaves evidence. Can you tell which version is serving traffic? Can you halt a rollout that looks wrong? Can you see the budget move during the release, not the next morning? You pair with the people who write the feature so the path is usable on a normal day, and you keep the rollback as rehearsed as the rollout. Handoffs between rotations are part of the same path. The note you leave names the budget remaining, the deploy still in progress, and the mitigation that is temporary. The next person should be able to continue your incident without reconstructing it from chat fragments. Capacity, saturation, and a dependency diagram belong in the same habit. A service that is reliable only when nothing changes is a service that has not met its own roadmap.

Proof is the service you have kept up

No licence makes someone a site reliability engineer. Employers look for a service you have taken responsibility for: an on-call rotation you completed, an incident write-up with a timeline, an error budget you operated, or a piece of toil you removed and can demonstrate. A computer science degree or a move from feature work, systems administration, or operations can all be the start. Cloud vendor credentials sometimes appear on resumes. They show familiarity with a provider's console. The hiring manager still asks what paged you and what you changed afterward. Preparation is supervised practice on a real rotation, a lab you can break safely, and the habit of writing down what the user experienced.

If you are early, build a small service, give it a health check, a dashboard, and a deliberate failure, then write the note you would have wanted at the start of the page. Put that note where a reviewer can read it. If your production stories are confidential, prepare the shape: the symptom, the signal that found it, the mitigation, and the lasting repair, with customer names removed. Be ready to whiteboard a rollout and a rollback. The proof employers trust is specific and a little boring. Hero stories without a follow-up change read as luck.

Bring one incident and one piece of toil

Walk in with a timeline you wrote and a manual chore you deleted. Say what the service does for a user, what "bad" looks like, and who is allowed to halt a release.

Joining a rotation that already pages someone

Postings in this work mention on-call, a named service, and words like availability, latency, or incident review. Read them for scope. One service with a team is a different seat from a platform that every product shares. Your application should name a service, your role when it failed, and the change that stuck. "Improved reliability" is a slogan. "Cut a noisy page, added a rollback the feature team could run, and wrote the budget rule the product owner accepted" is a record. Include the tools only after the story, so the reader meets the service before the vendor list.

Interviews often replay an incident. They want your sequence: what you checked first, what you ignored, when you communicated, and what you refused to change in the heat of the event. A second conversation may be a design: a new dependency, a rollout plan, a page that should exist. Ask how often the rotation fires, what happened after the last severe incident, and whether error budgets change the release calendar or sit in a slide. Ask who writes the follow-ups and who is allowed to miss them. Ask whether you will still change code. A reliability seat that cannot touch the service becomes a reporting function. A seat that only builds pipelines and never holds a budget is a different kind of work than the one this role describes.

Say plainly how you handle sleep and pages, without turning the interview into a contest. Ask whether follow-the-sun coverage exists, and what the team does after a hard night. Tell them constraints on location and sponsorship early. If you are moving from feature development, say so, and show one production problem you owned after your own release. That story lands better than a claim that you have always thought about systems in the abstract. The team is hiring a person who will be awake, accurate, and willing to stop a ship when the budget says stop.

From the person who answers to the person who sets the budget

The first rotation teaches you the service's actual failure modes, which are rarely the ones in the diagram. You learn the deploy tool, the noisy alerts, and the people who answer when you ask. The next step is owning reliability for that service: the budget, the weekly review of toil, and the list of repairs that feature work must make room for. Later, some engineers carry that practice across several services, helping other teams write targets they will honor. Some become the incident lead the company wants in the room. Some stay deep on one platform, capacity, or the release path, and that depth is a destination when the service is critical.

Evidence for a wider role is a budget people follow, an incident review that changed a roadmap, and toil that stayed gone. Titles differ. Describe whether you can halt a release, whether other teams adopt your review habit, and whether you still know the current failure modes or only the policy. A move into people management means the rotation health, hiring, and the argument with product leadership about how much unreliability the business will buy. Take that when you want the argument. Stay on the tools when the service still needs a person who can roll it back. Both paths are real, and a company that treats the individual path as a stall will tell you so in how it pays and how it assigns the next incident.

Keep a private log of incidents and of toil you removed, with dates and outcomes, confidential details stripped. That log becomes your promotion packet and your next interview. The growth in this career is a larger surface of production you can describe honestly: more services, harder dependencies, or a practice other teams copy because it made their pages rarer. The center stays the running service, the budget, and the path a change takes once it leaves a branch.

What an on-call offer is worth on this chart

A site reliability engineer who owns error budgets and the production path can measure an on-call offer with the May 2025 Software Developers wages from the Bureau of Labor Statistics Occupational Employment and Wage Statistics. The published entry figure is $82,460. The national median is $135,980. The gap between those two published figures is $53,520. A first rotation, still learning one service with a mentor on the page, can sit nearer entry. Ownership of the budget, the authority to halt a rollout, and a record of incident repairs are the scope that supports a conversation toward the median. If the offer includes nights and the number remains near $82,460, the $53,520 is the published distance you can cite, tied to the service you will be trusted to stop.

California's high end of the published range is $272,670, for places with enough people in the occupation for the Bureau to publish it. The chart puts $136,690 between the national median and that high end. California's median, typical pay in the state, is $174,410, and the gap from the national median to that state median is $38,430. Use $174,410 when you mean a typical California offer for this series. Use $272,670 when the role sets reliability practice across services and the company is paying at the far end of the published range. Those are different claims. An on-call rotation by itself does not pick the line. The scope does.

New York's median is $166,180. Washington's is $166,540. Massachusetts lists $165,210. Oregon lists $142,720. If two offers land in that New York and Washington pair, the Bureau medians are close, so compare the severity of the rotation, the staffing of follow-ups, and whether you can change the service you are paged for. Puerto Rico's median is $79,380, the lowest on the chart, and it sits below the national entry figure. Read a Puerto Rico offer against $79,380 and $82,460, and keep California's range high end in California. End the negotiation able to name the error budget you will watch and the single figure on the chart that matches the scope they described.

The top of Site Reliability Engineer pay — and how to get there with AI

$272,670what Site Reliability Engineer pay reaches in California

Highest state-level top-of-range annual wage for Software Developers, among states with at least 500 people in the job. U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025.

And the role it leads to — Computer Hardware Engineers — reaches $281,210 in California.

$82,460entry$135,980middle$272,670top end

A site reliability engineer reaches the top of this range by converting hours they removed from the group's week into formal ownership of something larger, rather than letting the saved time silently fill up again.

Reliability work has a cruel property: when it succeeds, nothing happens, so nobody sees it. Monitoring whether equipment and services function in conformance with their specifications, storing and manipulating data to analyse system capability and requirements, and preparing the reports that describe project status are all in the job description, yet the engineer who does them well often has the weakest promotion case in the room. The fix is an hours ledger. Count what repetitive work you deleted, in hours per week, then spend that credibility asking for the next thing, capacity forecasting, hardware configuration decisions, or supervising the technicians who carry the rota. Assistants make the deletion cheap; the counting is what makes it visible.

Your playbook, by where you are now

Just startingCount the repetitive work, then delete it

  1. Spend two weeks writing down every interruption and manual step, with minutes beside each, before you automate anything.
  2. Pick the three that cost the most hours and remove them, not the three that are most interesting.
  3. Write in Cursor or with GitHub Copilot for the plumbing, and keep the review standard the same as for anything else you ship.
  4. Hold the ledger in Airtable so saved hours per week can be shown as a line rather than described from memory.
  5. Take one service and write down what conformance to specification means for it, in a page a new engineer could use.

What proves it: A ledger showing hours per week removed, with the change that removed each one.

Realistic span: the first two quarters

A few years inTurn hours back into capability

  1. Spend the recovered time deliberately on one capability the group lacks, rather than on whatever pages loudest.
  2. Pull operational and cost data into Alteryx software so questions about system capability get answered from data instead of debate.
  3. Model growth against Amazon Elastic Compute Cloud EC2 and Amazon DynamoDB usage and publish where the top of the range actually sits.
  4. Write the quarterly report on project specifications, activity and status yourself, and make it the version other leads copy.
  5. Train the engineers and technicians who use anything you built, since a tool nobody else can operate is a liability with your name on it.

What proves it: A published capacity forecast that a purchasing or scaling decision was made from.

Realistic span: years two through four

ExperiencedAsk for the ground, with the ledger attached

  1. Obtain and evaluate the information on costs, security needs and reporting formats that determines hardware configuration, and own that call.
  2. Specify power, cooling and environmental requirements early enough that facilities and finance plan around them rather than react.
  3. Supervise and assign work to the programmers, technologists and technicians who keep the systems running, and grade handover quality alongside fixes.
  4. Recommend one purchase or one retirement each quarter with the before and after figures already gathered.
  5. Look at where this work prices highest, with California at the top, and treat hardware engineering as the widening move when the physical layer becomes the interesting problem.

What proves it: Written ownership of capacity and configuration decisions, with reporting that survives audit.

Realistic span: four years and beyond

The next 90 days

Before writing a line of automation, spend fourteen days keeping an honest ledger. Every ticket, every page, every manual step, every time somebody interrupted you to run something, with the minutes it took. Most engineers are startled by the total and more startled by where it concentrates, because it is almost never the thing they assumed. Delete the top item properly, then keep the ledger running so the saving shows up as a line instead of a claim. At the end of the quarter, take one page to your manager: here is what the group was spending, here is what it spends now, and here is the area I want to own next. That is a promotion conversation with evidence in it, which is rarer than good reliability work.

Wage figures: BLS OEWS, May 2025. The playbook is PayCrunch editorial guidance, not a guarantee of pay or placement.

Careers related to Site Reliability Engineer

Similar pay, same field

Where this can lead

Every figure is the national median from the U.S. Bureau of Labor Statistics (OEWS) shown on that role’s own page.

Never used AI before? Start here (2 minutes).

Turn on the AI already inside your observability stack. On Datadog enable Bits AI; on PagerDuty turn on its AI incident features; on Honeycomb or Grafana try the natural-language query assistants. Ask a real question about a recent incident and watch it correlate signals you'd have chased by hand. Zero new tools required.

For runbooks, postmortems, scripts, and config, keep Claude or ChatGPT open in a second tab (redact secrets and PII). It drafts the automation and the write-up; you verify and own every command that touches production. AI is the calm senior on-call who has already read all the dashboards.

The one rule, forever: Never let AI execute a remediation in production without a human in the loop and a rollback plan — an AI-suggested 'fix' during an incident can widen the outage, and root-cause guesses are often wrong. Verify before you act. And never paste production secrets, customer data, or raw logs containing PII into a consumer AI tool; use observability and incident platforms your org has vetted.
The plays — exact steps, exact prompts

Do these in order. Each one is copy-paste ready. You do not need to know anything about AI going in.

1
Cut MTTR with an AI incident copilot
Why this pays: Mean-time-to-resolution is the metric SREs live by. AI that correlates alerts, drafts the timeline, and suggests likely causes shrinks MTTR — the reliability outcome that earns senior SRE pay and calmer nights.
PagerDuty (AIOps)Rootlyincident.ioDatadog Bits AI
1
Let your incident platform (PagerDuty, Rootly, incident.io) auto-assemble the timeline, correlate alerts, and suggest probable causes — then verify every suggestion before acting.
2
Work an active incident with this prompt.
Copy-paste this prompt
You are an experienced SRE helping me during an incident. Symptoms: [API p99 latency jumped from 200ms to 4s at 14:32; error rate up; a deploy went out at 14:20]. Recent changes: [paste]. Give me a prioritized hypothesis list from most to least likely, the exact read-only checks to confirm or rule out each (which metrics, logs, traces), and for the top hypothesis the safest mitigation and how to roll it back. Flag anything destructive.
AI root-cause guesses are hypotheses, not answers — confirm with real telemetry before you change anything, and keep a rollback ready.
What you'll haveFaster, calmer incident resolution with a shorter MTTR — the reliability outcome that earns senior SRE pay and far better nights on call.
2
Write blameless postmortems and runbooks in minutes
Why this pays: Good postmortems and runbooks prevent repeat outages, but writing them is the chore SREs skip. AI drafts them from the incident data, so the learning actually gets captured — the discipline that compounds into reliability and seniority.
Rootlyincident.ioClaudeNotion
1
Generate the postmortem draft from the incident timeline and your notes, then add the human judgment — contributing factors and real, owned action items — yourself.
2
Draft the postmortem with this prompt.
Copy-paste this prompt
Act as an SRE writing a blameless postmortem. Here is the incident data: [paste timeline, impact, what was done]. Draft a postmortem with: summary, customer impact, timeline, root cause and contributing factors, what went well, what went poorly, and specific, owned action items with the reliability improvement each delivers. Keep it blameless — focus on systems, not people.
Keep it blameless and make the action items real and owned — an AI draft full of vague 'improve monitoring' items helps no one. You supply the judgment.
What you'll havePostmortems and runbooks that actually get written and capture the lesson — the compounding reliability discipline that separates a senior SRE from a firefighter.
3
Query logs, metrics, and traces in plain English
Why this pays: Debugging speed depends on getting answers out of telemetry fast. Natural-language querying lets you interrogate logs, metrics, and traces without memorizing every query language — so you find the needle faster and take on the hairiest, highest-value systems.
Datadog Bits AIHoneycombGrafanaNew Relic
1
Use your platform's AI query assistant (Datadog Bits AI, Honeycomb, Grafana) to ask questions in English, then confirm the generated query actually matches what you meant.
2
Build a query with this prompt.
Copy-paste this prompt
Help me write a [Datadog / PromQL / Honeycomb] query. I want to see [the p99 latency of the checkout service broken down by pod and endpoint over the last 6 hours, marking the moment error rate crossed 1%]. Explain what the query does field by field, and suggest two related views that would help me tell a code problem apart from an infrastructure problem.
Read the generated query before you trust its output — a subtly wrong filter or aggregation can point you at the wrong service during an incident.
What you'll haveAnswers pulled from telemetry in seconds without wrestling query languages — the debugging speed that lets you own the most complex, highest-value systems.
4
Engineer away toil with AI automation
Why this pays: Toil — repetitive manual ops — is what burns SREs out and causes mistakes. AI writes the scripts, playbooks, and IaC to automate it, freeing your time for real reliability engineering, which is what the top of the band actually does.
AnsibleTerraformClaude CodePython
1
Pick a recurring manual task and have AI draft the automation, then test in staging with dry-run and rollback guardrails before it touches production.
2
Automate a toil task with this prompt.
Copy-paste this prompt
Act as an SRE automating toil. I currently [manually rotate and validate TLS certs across 20 services once a quarter]. Write [an Ansible playbook / Python script] to automate it safely: idempotent, with a dry-run mode, pre-checks, clear logging, and a rollback path. Explain each safety mechanism, and list what I should test in staging before running it in production.
Automation that fails badly is worse than manual toil — insist on dry-run, idempotency, and a rollback, and prove it in staging first.
What you'll haveRecurring toil turned into safe, tested automation — the reclaimed engineering time that lets you build real reliability, which is the work top-of-band SREs are paid for.
5
Design SLOs, error budgets, and capacity plans with AI
Why this pays: Setting the right reliability targets — SLOs and error budgets — is the strategic core of SRE, and capacity planning keeps systems from toppling. AI helps you model both rigorously, the higher-order work that distinguishes a senior SRE.
ClaudeGrafanaPrometheusNobl9
1
Use AI to structure SLO definitions and error-budget policy, and to model capacity from your growth and load data — then verify the assumptions with your team.
2
Define SLOs with this prompt.
Copy-paste this prompt
Act as a reliability architect. Help me define SLOs for [a checkout API]. Given [current traffic, latency distribution, and the business tolerance for downtime], propose: the right SLIs to measure, target SLO levels with the reasoning, an error-budget policy (what happens when it's burned), and how to alert on burn rate rather than raw thresholds. Explain the trade-offs of tighter versus looser targets.
SLO targets are a business decision as much as a technical one — use AI to structure the options, then set them with your team and product owners.
What you'll haveWell-reasoned SLOs, error-budget policy, and capacity plans — the strategic reliability work that marks a senior SRE and carries the pay to match.
6
Master AI diagnostics and lead your team's AIOps
Why this pays: The SRE who masters AI-driven diagnostics and leads the team's AIOps adoption becomes the reliability authority — the route to staff or principal SRE pay at the top of the band.
k8sgptDatadogDynatrace (Davis AI)Claude
1
Master k8sgpt and your platform's AIOps (Datadog Watchdog, Dynatrace Davis) for anomaly detection and root cause, then pilot the approach for your team with measured MTTR results.
2
Propose the AIOps pilot with this prompt.
Copy-paste this prompt
Create a 1-page proposal to adopt AIOps on our team. We run [Kubernetes on AWS with Datadog]. Cover: where AI-assisted anomaly detection and incident correlation fit our current on-call workflow, the metrics to prove impact (MTTR, alert noise, on-call load), the risks (over-trust, alert fatigue, false root causes) and how to manage them, and a 30-day pilot on one service. Audience: my engineering manager.
Prove it with MTTR and alert-noise numbers from a real pilot — that is what earns the mandate and the staff-level title.
What you'll haveA team-wide AIOps practice you lead with measured results — the reliability authority that carries staff-level pay near $272,670.
Your 12-month sequence to the top of the range

How the plays above stack into a path from median pay toward the $272,670 tier.

Month 1
Turn on the AI in your observability and incident stack (Datadog Bits AI, PagerDuty, Honeycomb). Use it on one real question and one recent incident.
Months 2-3
Put an AI incident copilot to work on-call — verifying every hypothesis — and draft your next postmortem with AI.
Months 3-6
Automate one painful source of toil with AI-written scripts or IaC, tested in staging with dry-run and rollback.
Months 6-9
Use AI to define SLOs and error budgets for a key service and model its capacity.
Months 9-12
Master AI-driven diagnostics (k8sgpt, Watchdog/Davis) and run an AIOps pilot with MTTR metrics.
Year 2
Lead your team's AIOps adoption and reliability strategy — the staff/principal route to the $272,670 tier.
Gear for this job

As an Amazon Associate, PayCrunch earns from qualifying purchases. Links to books and tools are for the job on this page; we only recommend what we’d use in the work.

Burns / Beda / Hightower Kubernetes: Up and Running, 3rd

Same live O’Reilly 3rd already on cloud-engineer. This page names Terraform and Kubernetes as the reliability stack and the AIOps proposal runs Kubernetes on AWS. Not Terraform Up and Running as the lead (that is the IaC book on devops-architect / automation-engineer) and not CompTIA Security+ (that is software-engineer / infosec).

Next steps for a Site Reliability Engineer

Some links below are affiliate or partner links. PayCrunch may earn a commission if you enroll or subscribe through them, at no extra cost to you. Wage figures on this page still come from the Bureau of Labor Statistics, not from these programs.

Site Reliability Engineer work is specific enough that a stamped 'check out these courses' block would be noise. BLS files this work as Software Developers (SOC 15-1252). O*NET Job Zone 4 is typical: a bachelor's degree, so the honest next credential is a professional certificate or bachelor's-level coursework — not a random catalog dump.

Site Reliability Engineers in this dataset list AJAX among the tools in use, so a program that names that stack is a better fit than a survey course.

The next title this dataset points at is Computer Hardware Engineers; a credential aimed that way is a clearer step than another year in the same seat.

Computer Science programs on Coursera for Site Reliability Engineer work

Coursera search for computer science — a professional certificate or bachelor's-level coursework that lines up with computing, not a generic professional-development aisle.

Computer Science courses on edX

edX search for computer science, aimed at computing (SOC 15-1252). Same field as the Coursera link, different university catalog.

Screened remote and flexible Site Reliability Engineer listings on FlexJobs

FlexJobs screens remote, hybrid, freelance, and flexible listings so you are not wading through unverified ads. This is a job-board search for Site Reliability Engineer work, not a claim that they list a counted SOC 15-1252 inventory.

Build a Site Reliability Engineer resume on Resume Now

Write a Site Reliability Engineer resume, or one aimed at Computer Hardware Engineers, instead of a blank template. Resume Now is a resume builder; we are not claiming a counted template set for this SOC.

Build a Site Reliability Engineer resume on Zety

A Site Reliability Engineer resume that names the actual tasks on this page, or the step-up title Computer Hardware Engineers, beats a blank template when you apply.

What Site Reliability Engineers earn by state

These are the Bureau of Labor Statistics’ own figures for Software Developers, state by state — not a cost-of-living adjustment applied to the national number. Only states employing at least 500 people in the occupation are shown, because a state median drawn from a handful of workers is noise rather than a signal.

California
$174,410
highest of them · +28% vs the national median
Puerto Rico
$79,380
lowest of the 51 states and territories that qualify · -42% vs the national median
The same job pays $95,030 more a year at the median in California than in Puerto Rico — 120% higher. That gap is what the Bureau measured, before any question of what it costs to live in either place. California also carries the top of this job’s range, $272,670 — the figure quoted at the head of this page.
California$174,410Washington$166,540New York$166,180Massachusetts$165,210Oregon$142,720New Hampshire$139,720Maryland$138,680Colorado$138,390

Source: U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025, SOC 15-1252. 51 states and territories clear the 500-employee reporting floor for this occupation; those below it are left out rather than shown with a wide error band.

Free data. Use any of it.

PayCrunch publishes verified, BLS-sourced salary + AI-playbook data on 1,000+ professions — free, no signup.

Frequently asked
Will AI replace site reliability engineers?
No — it augments them. AI correlates signals and drafts the write-up, but deciding what to do under pressure, owning the production change, and engineering systems to be reliable in the first place are human. AI makes SREs faster and frees them from toil; it doesn't carry the pager's accountability.
Can I trust AI's root-cause suggestion during an incident?
Treat it as a ranked hypothesis, never a verdict. AI correlates and guesses; it is often wrong, and an AI-suggested fix can widen an outage. Confirm with real telemetry and keep a rollback ready before you change anything in production. You own the outcome.
Is it safe to use AI with our telemetry and incidents?
Only within vetted tools. Never paste production secrets, customer data, or PII-laden logs into consumer AI. Use the AI built into your observability and incident platforms, which run inside your security boundary, and redact anything sensitive from general tools.
How does AI actually raise an SRE's pay?
By improving the metrics SREs are judged on. Lower MTTR, less alert noise, less toil, and better SLOs mean more reliable systems and more time for real reliability engineering — the outcomes that earn senior and staff SRE comp toward the top of the band.
What's the difference between an SRE and a platform engineer using AI?
SREs own reliability, on-call, and incident response; platform engineers build the self-service developer platform. AI helps SREs most with incident response, postmortems, and telemetry querying, while platform engineers lean on it for scaffolding and infrastructure-as-code. The roles overlap but the focus differs.
Methodology & sources
  • Salary (median, 10th, top of the range) — U.S. Bureau of Labor Statistics, OEWS.
  • By state — the Bureau of Labor Statistics’ own state medians, limited to states employing at least 500 people in the occupation. No cost-of-living arithmetic is applied to a wage anywhere on this page.
  • The plays — PayCrunch's own step-by-step guidance using publicly available AI tools. Tool names/URLs are real and current as of August 2026; prompts written to work as-is. Verify any professional output before relying on it.

Sources