PayCrunch Research · The exact AI playbook for your profession, sourced to the U.S. Bureau of Labor Statistics

PayCrunch AI Playbook · Technology

What lifts an ethics researcher to the top of the range

$282,110estimated top of the range · middle $125,000 / yr
AI augments this role

AI Ethics Researchers in the United States earn a median of $125,000 a year. Pay starts near $78,000. The top of the range is estimated at $282,110. The Bureau of Labor Statistics does not publish a separate wage series for this exact title, so this figure is derived from the closest occupation it does track and is labelled an estimate.

Source: PayCrunch estimate. Last checked 9 September 2026.

Entry level
$78,000
Top-end estimate
$282,110
Education
Master's or Doctoral degree
Lower disruption Higher exposure AI augments this role
Entry · $78,000 Top-end estimate · $282,110 Middle $125,000

Wages — PayCrunch estimate. The Bureau of Labor Statistics does not publish a separate wage series for AI Ethics Researcher; figures are derived from the closest occupation it does track and are labelled as estimates. AI-impact rating is PayCrunch's editorial assessment. Updated September 2026.

🆕 New & Trending AI Tools for AI Ethics ResearcherReviewed September 2026

We track new AI-tool launches every week and refresh this list — here’s what’s gaining traction for AI Ethics Researcher work right now.

Julius AINEWFree / $20 mo

AI data analyst that runs statistics and charts from plain-language prompts.

How an AI Ethics Researcher uses it: analyze datasets and generate figures without writing code

NotebookLMNEWFree / $7.99 mo

Google tool that answers questions grounded only in the documents you give it — with citations.

How an AI Ethics Researcher uses it: load your own manuals, policies, or PDFs and ask questions that stay accurate to the source

ElicitFree / $12 mo

AI research assistant that finds and summarizes papers.

How an AI Ethics Researcher uses it: run a literature review and extract findings across dozens of papers fast

ConsensusFree / $9 mo

AI search that answers questions from peer-reviewed research.

How an AI Ethics Researcher uses it: get evidence-backed answers with the studies behind them

SciSpaceFree / paid

AI that explains papers and helps with literature review.

How an AI Ethics Researcher uses it: decode dense papers and trace citations quickly

SciteFree / $20 mo

Shows whether other studies support or contradict a paper's claims (Smart Citations).

How an AI Ethics Researcher uses it: check if a finding is actually backed by the wider literature before you cite it

ChatGPTFree / $20 mo

The most-used AI assistant — writing, analysis, research, and images from a plain-language chat.

How an AI Ethics Researcher uses it: draft emails and documents, summarize long files, and get instant answers to on-the-job questions

ClaudeFree / $20 mo

AI assistant known for careful writing, long-document analysis, and coding.

How an AI Ethics Researcher uses it: analyze big reports or spreadsheets and turn messy notes into clean, finished writing

Google GeminiFree / $20 mo

Google's AI assistant, built into Gmail, Docs, and Search.

How an AI Ethics Researcher uses it: draft and reply inside Google Workspace and research without leaving the page

Policy memos, a philosophy seminar, legal research, or social-science fieldwork may already be your craft. An AI ethics research role asks you to aim that craft at automated systems and at the people who have to live with their decisions. This letter is about the research week, the proof that gets you hired when no licence exists, the path after the first research seat, and the pay figures on this page, which are estimates and need to be described that way.

Research on what automated systems do to people

The day is reading, writing, and evidence. You might study a hiring model, a benefit-eligibility system, a medical triage tool, a content ranker, or a scoring system a city uses on residents. You ask who is affected, which errors fall on whom, who can appeal, and who is accountable when the system is wrong. You read the documentation the builder will share, the complaints users have already made, and the prior studies in the same domain. You design an inquiry that can be repeated: a record review, an interview study, a comparison of outcomes across groups, or a structured evaluation of a model’s behavior on cases you can describe. Then you write so a lab, a company, a regulator, or a journal can see the claim and the limit of the claim.

Fairness in this job is a research object, not a slogan. You define the harm you are looking for, the population, and the comparison, and you say what your method cannot see. Accountability is similarly concrete. You trace which team can change the system, which notice a person receives, and which record exists if someone challenges a decision. Coming from law, you already know how to read a duty. Coming from sociology or public policy, you already know how institutions distribute burden. Coming from philosophy, you already know how to pressure a definition. The new muscle is attaching that judgment to a real system’s inputs, outputs, and owners, and showing your work to people who build models for a living.

The tools are a literature library, a notes system, spreadsheets or a statistics environment when the study is quantitative, and the collaboration space of the lab or the company. The places are a university group, a corporate research team, a nonprofit, a government office, or a lab inside a foundation. The people are engineers, product managers, affected communities, lawyers, and the principal investigator or research director who decides which study is worth the quarter. Your decision is often what not to claim. A vivid anecdote can illustrate. It cannot carry a finding about a whole system. Learning to mark that line is the craft.

Some roles lean empirical and sit close to evaluation: sampling cases, comparing error patterns, documenting where a model’s behavior departs from a stated purpose. Some lean normative and sit close to governance: proposing the rule a deployer should adopt, the appeal a person should have, the record that should exist. Strong researchers can visit both rooms. Early in a career change, pick the room your previous field already trained. A statistician who forces a philosophical essay will look thin. A lawyer who invents a benchmark will look thin. Adjacent excellence, pointed at systems, is the credible offer.

A finished study has a shape you can reuse. State the system and the decision it influences. State whose outcomes you examined and whose you could not see. State the alternative explanation, including the possibility that the harm sits in the policy around the model rather than in the model alone. State what you would need before recommending a pause, a redesign, or a narrower use. That last sentence is where career changers from advocacy sometimes overreach, and where career changers from engineering sometimes go quiet. Practice writing the recommendation at the strength the evidence supports, then stop. Reviewers in this field trust a bounded claim more than a sweeping one.

Writing, a research degree, or policy work

No licence, three kinds of proof

No agency issues a licence to be an AI ethics researcher. Hiring managers look for writing that shows a finished argument, a research degree with a study you can defend, or policy work where you influenced a real decision about an automated system.

Writing means a paper, a public report, a well-argued memo, or a documented study, not a thread of opinions. For each piece, a reader should see the system, the people affected, the method, and the limit. A research degree, often a master’s or a doctorate in a social science, law, philosophy, information studies, computer science, or public policy, proves you can finish a long inquiry under supervision. Policy work proves you have sat where rules get written: a comment on a proposed rule, a procurement clause you helped draft, an internal review that stopped or changed a deployment. Any one of the three can open a door. Two of them make the door ordinary rather than exceptional.

Preparation follows the proof you are building. If you need the degree, choose a program where you can study real systems with a supervisor, and leave with a document you may share. If you already have the degree in an adjacent field, add one study about an automated system before you retitle yourself. If you are in government or advocacy, write one public piece that shows how you handled evidence, uncertainty, and remedies. Course certificates that promise a short tour of ethics vocabulary are optional reading, not the credential. Employers can tell a glossary from a finding.

Collaboration with technical people is part of preparation too. You do not have to train models to research their effects, and you do have to read enough of an evaluation report to know when a metric is being asked to stand in for a moral conclusion. Sit with an engineer long enough to learn how a dataset was built and who was left out of it. That literacy is what keeps your writing from floating above the system.

How a lab or an office actually hires

University groups hire through advisors and postings for research assistants, postdoctoral fellows, and faculty. Companies hire researchers into responsible-innovation teams, research labs, and sometimes into policy groups that sit beside product. Nonprofits and government hire people who can turn a docket of systems into findings a decision-maker will read. Apply to the setting whose week you want. A faculty path is writing, teaching, and funding. A company path is research that arrives in time to change a launch. A government path is research that can survive a record and a public challenge. Those clocks differ, and your samples should match the clock.

The interview is a discussion of one study. Walk through the choice of cases, the alternative explanation you considered, and the recommendation you were willing to make. Bring the document. If the study is confidential, bring a rewritten version that keeps the method and drops the client. Expect a conversation about disagreement. Someone in the room may think your harm is overstated or understated. The skill is to revise the claim in public without abandoning the evidence. People who treat pushback as an insult struggle in this job, because the job is pushback, in both directions.

Coming from journalism, emphasize verification and the correction you issued. Coming from law, emphasize the remedy and the limit of what a rule can do. Coming from engineering, emphasize a system you slowed down because the evaluation was thin, and show the write-up, not only the instinct. Coming from community organizing, emphasize how you gathered accounts without exposing people to harm, and how you separated a pattern from a single story. Each of those origins is legitimate. Each needs a finished artifact.

Researcher, then a lead on the program

After the first research seat, the path runs toward a senior researcher and then a lead who sets the program: which systems get studied, which methods the group will stand behind, and how findings reach the people who can change the deployment. In a university that lead is often a principal investigator. In a company it may be a research manager or a head of responsible innovation who still reads the work. In government it may be a senior advisor with a small staff. Some people step sideways into policy roles that use the research rather than produce the next paper. That step is a real career. It should be chosen because you want the decision room, not because the writing got hard.

What earns the wider seat is a body of work other people can use. A single viral essay fades. A sequence of studies, memos, or evaluations that changed a design, a rule, or a research agenda is the promotion case. Keep a private list of those outcomes, careful with confidential details, so you can tell the story later without reconstructing it from memory. Mentor one junior person along the way. Leads are hired for judgment and for the ability to multiply it.

If you want to stay individual-contributor and go deep, say that early. This field sometimes treats management as the only raise. A staff-level researcher who owns a method, a domain, or a long partnership with affected communities is a different and valid peak. Ask whether the organization has that peak before you assume the only up is people-management.

The wages here are estimates, full stop

Entry on this page is $78,000. The median is $125,000. The high end is $282,110. These three figures are estimates. The Bureau of Labor Statistics does not publish a separate wage series for the exact title AI ethics researcher, so the page derives them from the closest occupation it does track and labels them estimates. Call them estimates in every conversation. Reserve the Occupational Employment and Wage Statistics label for titles the Bureau publishes under their own names. Keep the quotes national. Nothing on this page attaches $78,000, $125,000, or $282,110 to a state.

The gap from entry to the median is $47,000. The gap from the median to the estimated high end is $157,110. Use those gaps only as distances inside an estimate. An offer near $78,000 sits at the entry end of this estimate. An offer near $125,000 sits at the middle. The $47,000 between them is a reasonable way to describe moving from a first research seat, or from an adjacent policy wage that mapped to the bottom, toward the middle of the estimated range. The $157,110 above the median describes the room the estimate allows for senior research leadership. Speak of $282,110 as the top of this page’s estimated range.

This page also shows a May 2025 employment count of 37,200. Read that count as the nearest occupation the Bureau tracks, the figure placed alongside the estimate. It belongs with that neighboring occupation. Leave 37,200 out of the salary sentence. A headcount for the nearest series is context for how the page ordered the title, and the offer conversation stays with the three estimated wages and the two gaps.

A clean way to talk: “The public estimate I am looking at puts entry at $78,000 and the middle at $125,000, with an estimated high end of $282,110. I am anchoring on the middle because the role owns studies, and I am describing the high end as an estimate.” Then connect the anchor to your proof: a finished study, a degree, or policy work that changed a system. If the employer has internal bands, ask where those bands sit relative to $125,000. If they cite a city premium, ask them to justify it from their own budget, since a state median for this title is absent here. Senior scope, a scarce domain, or a role that leads other researchers is the honest reason to gesture at the $157,110 above the median, and the destination stays an estimate when you do.

What to keep from the field you are leaving

Keep the method your old field already respects, and point it at a system with owners and people on the receiving end. Retire the urge to retitle every opinion as research. You get hired with writing, a research degree, or policy work, and you grow by accumulating findings other people can use. When money comes up, stay inside $78,000, $125,000, and $282,110, name them as estimates, and let the $47,000 and $157,110 gaps describe scope inside that estimate. That honesty is part of the job you are asking to do.

The top of AI Ethics Researcher pay — and how to get there with AI

$282,110top-end estimate for AI Ethics Researcher

PayCrunch estimate - derived from the closest occupation BLS tracks (Computer and Information Research Scientists, 15-1221). This figure is PayCrunch’s estimate, not a Bureau of Labor Statistics published wage for this exact title.

And the role it leads to — Natural Sciences Managers — reaches $330,050 in California.

$78,000entry$125,000middle$282,110top end

The middle of this range writes position papers that circulate internally; the top of the range belongs to the researcher whose evaluation evidence a customer's counsel accepts before a contract is signed.

Two researchers can hold identical views about model harm and be paid very differently. The difference is usually who is in the room when a buyer asks whether the system can be trusted. Developing and interpreting organizational policies and procedures is internal work; consulting with users, management and vendors to pin down requirements, then evaluating project plans and proposals for feasibility, is the same skill aimed at a deal. Produce the evidence — a rerunnable evaluation, a performance standard teams are measured against, a written answer to what procurement always asks — and the sales engineers stop being able to close without you.

Your playbook, by where you are now

Just startingMake one failure mode measurable

  1. Choose a single behaviour your team argues about and define it as a test with stated inputs, expected behaviour and a clear miss condition.
  2. Assemble the first case set by hand, small enough that you can defend every item, and keep results beside the model outputs in Amazon Redshift so runs stay comparable.
  3. Version the suite in Apache Subversion SVN next to the code, so a regression is demonstrable rather than debatable.
  4. Put Claude or Gemini on the adversary side: ask it for the inputs most likely to break your stated rule, then hand-check everything before it enters the suite.

What proves it: An evaluation suite engineering runs before a release rather than after one.

Realistic span: the first year on the team

A few years inTurn the evaluation into a standard

  1. Write the performance standard: what thresholds gate a release, who signs, and what happens on a failed run.
  2. Move the suite onto Apache Spark over Amazon Web Services AWS software so it runs against a full sample instead of a convenient slice.
  3. Join a multidisciplinary build — human-computer interaction, robotics, virtual reality — early enough to change the interface rather than review it.
  4. Keep an answer library holding every question auditors, procurement or a customer's counsel has raised, each with your evidence attached.
  5. Take the technical trust questions yourself on customer calls instead of briefing someone else to relay them.

What proves it: A signed release standard plus an answer library the field team reuses without editing.

Realistic span: years two through five

ExperiencedBe why the agreement closes

  1. Lead the assurance review on the largest opportunities: scope it, run it, and present findings to the customer in person.
  2. Publish the method openly enough that another team can repeat it — data, metric, threshold, and the holes you already know about.
  3. Assess proposals for feasibility before anything is promised to a client, and record plainly what the system cannot yet do.
  4. Stand up the internal review board, set who may release what, and measure teams against those rules on a fixed cycle.

What proves it: A closed enterprise agreement whose assurance section you wrote and defended.

Realistic span: seven years onward

The next 90 days

Over the next quarter, take one claim your organization makes about its system — that it refuses a category of request, that it treats two groups alike, that it stays inside a stated limit — and prove or disprove it with a test anyone could rerun. Build the cases yourself. Log every result against the model version and the date. Then write two pages: what you measured, what came back, and the honest caveat. Send it to the people who talk to customers, not only to the research channel. The first researcher who hands a field team defensible evidence instead of a warning stops being counted as overhead.

Wage figures: PayCrunch estimate. The playbook is PayCrunch editorial guidance, not a guarantee of pay or placement.

Careers related to AI Ethics Researcher

Similar pay, same field

Where this can lead

Every figure is the national median from the U.S. Bureau of Labor Statistics (OEWS) shown on that role’s own page.

Never used AI before? Start here (2 minutes).

Start by making your evaluations reproducible. Open the UK AI Safety Institute's Inspect framework (open-source) or Stanford's HELM and run a standard benchmark against the model you care about, so you have a versioned, repeatable baseline instead of vibes. That rigor is what separates a researcher from a commentator.

For the reading-and-synthesis half of the job, use Elicit or Perplexity to map the literature and NotebookLM to interrogate a stack of papers you have uploaded — always following citations back to the primary source before you rely on a claim.

The one rule, forever: Independence and rigor are the whole job. Never let the model you are auditing grade its own homework without disclosure; report AI-assisted findings as AI-assisted, not as verified fact; and make every eval reproducible (fixed seeds, versioned prompts, published methodology). Protect human-subject and red-team data, disclose your own AI use, and never launder a model's confident output into a citation.
The plays — exact steps, exact prompts

Do these in order. Each one is copy-paste ready. You do not need to know anything about AI going in.

1
Run reproducible model evaluations, not vibes
Why this pays: Rigorous, reproducible evals are the field's hard currency; the researchers who produce them (not opinions) get hired, cited, and promoted into lab and governance roles at the top of the band.
Inspect (UK AISI)Stanford HELMEleutherAI lm-eval-harness
1
Define the harm or capability you are measuring, then implement it as a versioned eval in Inspect or run it through HELM/lm-eval-harness. Fix seeds, pin model versions, and log every prompt.
2
Design the protocol before you build it, and pressure-test the methodology.
Copy-paste this prompt
I want to evaluate whether [model X] exhibits [sycophancy — changing a correct answer when the user pushes back]. Design a rigorous eval: operational definition, a dataset-construction plan, prompt templates, scoring rubric, controls to rule out confounds, and the statistics I would report. Note the threats to validity.
Use to design the protocol; you build and run it, and you review the methodology critically — the AI will miss confounds.
3
Publish the methodology and code so others can reproduce it — reproducibility is your reputation.
What you'll haveA portfolio of reproducible evals with your name on them — the credential that gets you into frontier-lab and governance roles near the top of the band.
2
Measure bias and fairness quantitatively
Why this pays: Responsible-AI hiring wants people who can quantify disparate impact, not just discuss it. Concrete fairness metrics on real systems are billable, publishable, and promotable.
FairlearnIBM AI Fairness 360SHAP
1
Use Fairlearn or AI Fairness 360 to compute group-fairness metrics (demographic parity, equalized odds) on a model's predictions, and SHAP to explain which features drive the disparities.
2
Frame the analysis and the tradeoffs before you run the numbers.
Copy-paste this prompt
For a [loan-approval classifier], list the fairness metrics I should report and their tradeoffs (why you cannot satisfy all simultaneously), how to choose the right one for [this regulatory context], and how to explain the tradeoff to non-technical stakeholders.
Use to frame the analysis and the write-up; run the actual metrics on the real model, and never overstate what a single metric proves.
3
Turn the numbers into a plain-language impact statement product and policy teams can act on.
What you'll haveFairness audits that hold up under scrutiny — the deliverable that makes you the go-to reviewer and lifts your pay.
3
Red-team models systematically
Why this pays: Red-teaming and jailbreak research is one of the hottest, best-paid specialties at frontier labs and safety orgs; systematic red-teamers are scarce.
GiskardGarak (LLM scanner)Claude & GPT (adversary and target)
1
Build an adversarial test suite in Giskard or the Garak LLM scanner, covering prompt injection, data exfiltration, harmful-content elicitation, and jailbreaks. Track attack-success rates across model versions.
2
Enumerate the attack surface systematically before you test.
Copy-paste this prompt
Act as a red-team lead. Given a chatbot that [answers customer billing questions and can call a refund tool], enumerate the top attack categories (prompt injection, tool misuse, PII exfiltration, jailbreaks), and for each give 3 concrete test cases and a pass/fail criterion. This is for authorized safety testing of my own system.
Only test systems you are authorized to test; log everything; report responsibly through disclosure channels, never publicly weaponize a live exploit.
3
Document findings with severity and reproduction steps and route them through responsible disclosure.
What you'll haveA track record of finding real vulnerabilities before they ship — exactly the profile frontier labs pay top-of-band for.
4
Synthesize the literature and the regulatory landscape faster
Why this pays: Whoever can map what is known — across ML, law, and policy — fastest becomes the person who writes the influential report or standard, which is where reputation and pay compound.
ElicitPerplexityNotebookLM
1
Use Elicit to build a structured literature matrix (claims, methods, evidence) and Perplexity for current regulatory developments, then load the key PDFs into NotebookLM to query them with citations.
2
Translate a regulation into concrete engineering obligations.
Copy-paste this prompt
Summarize how [the EU AI Act's high-risk requirements] map to concrete engineering obligations (documentation, logging, human oversight, robustness testing). Organize it as a checklist an ML team could actually implement, and flag where the text is ambiguous. Cite the specific articles.
Use to draft; verify every article and citation against the primary legal text — models routinely hallucinate section numbers.
3
Turn the synthesis into a briefing or standard proposal that engineers and policymakers can both use.
What you'll haveAuthoritative syntheses that get cited and shape decisions — the influence that defines a top-of-range researcher.
5
Ship model cards, datasheets, and governance artifacts
Why this pays: Operationalizing governance (NIST AI RMF, model cards, risk assessments) is what companies are staffing up to do for the EU AI Act era — the practical, well-paid side of the field.
NIST AI RMFModel Cards / datasheetsClaude / ChatGPT
1
Map a system to the NIST AI Risk Management Framework and generate the governance artifacts — model card, data datasheet, risk register — with AI drafting the boilerplate from your inputs.
Copy-paste this prompt
Draft a model card for [an internal resume-screening model]. Include intended use, out-of-scope uses, training-data provenance and known gaps, evaluation results by subgroup, ethical considerations, and a monitoring plan. Leave clearly-marked [TODO] placeholders wherever I must supply measured results.
AI drafts structure; you fill in real, measured numbers and sign off — a fabricated model card is a liability, not a shortcut.
2
Stand up a lightweight review process so these artifacts get produced for every launch, with you as the reviewer of record.
What you'll haveA repeatable governance program you own — indispensable to any company shipping AI under regulation, and a direct route to lead/principal pay.
6
Write the paper or policy brief that builds your name
Why this pays: Publications and influential briefs are how researchers build the reputation that commands frontier-lab and think-tank salaries; AI removes the drafting friction.
Claude / ChatGPTZoteroOpenReview
1
Use Claude or ChatGPT to pressure-test your argument (steelman the counter-position), tighten the structure, and turn results into clear prose — you supply every claim and citation.
Copy-paste this prompt
Here is my draft abstract and key results: [paste]. Play a skeptical FAccT reviewer: list the three strongest objections, the experiments a reviewer would demand, and where my claims outrun my evidence. Then suggest a tighter framing.
Use for critique and editing; never let it invent citations or results — verify everything against your data and Zotero library.
2
Submit to a venue that matters (FAccT, NeurIPS, an AISI report) and present it — visibility is the multiplier.
What you'll havePublished, cited work that establishes you as a voice in the field — the reputation that anchors top-of-band offers.
Your 12-month sequence to the top of the range

How the plays above stack into a path from median pay toward the $185,000 tier.

Month 1
Run one standard benchmark in Inspect or HELM to establish a reproducible baseline; start a literature matrix in Elicit.
Months 2-3
Add a quantitative fairness or robustness eval on a real system with Fairlearn/AIF360 and write it up.
Months 3-6
Build a systematic red-team suite and practice responsible disclosure on a system you own.
Months 6-9
Operationalize governance — map a system to NIST AI RMF and ship a model card and risk assessment.
Months 9-12
Turn your strongest result into a paper or policy brief and submit it to a real venue.
Year 2
Lead an eval or governance workstream and mentor others — the step into principal/lead roles.
Next steps for an AI Ethics Researcher

Some links below are affiliate or partner links. PayCrunch may earn a commission if you enroll or subscribe through them, at no extra cost to you. Wage figures on this page still come from the Bureau of Labor Statistics, not from these programs.

AI Ethics Researcher work is specific enough that a stamped 'check out these courses' block would be noise. BLS files this work as Computer and Information Research Scientists (SOC 15-1221). O*NET Job Zone 5 is typical: graduate or professional school, so the honest next credential is a graduate-level or professional certificate — not a random catalog dump.

The occupation's listed knowledge areas include Engineering and Technology and Design; the links search those subjects, not a generic 'career courses' list.

AI Ethics Researchers in this dataset list Amazon DynamoDB among the tools in use, so a program that names that stack is a better fit than a survey course.

Engineering And Technology programs on Coursera for AI Ethics Researcher work

Coursera search for engineering and technology — a graduate-level or professional certificate that lines up with computing, not a generic professional-development aisle.

Engineering And Technology courses on edX

edX search for engineering and technology, aimed at computing (SOC 15-1221). Same field as the Coursera link, different university catalog.

Screened remote and flexible AI Ethics Researcher listings on FlexJobs

FlexJobs screens remote, hybrid, freelance, and flexible listings so you are not wading through unverified ads. This is a job-board search for AI Ethics Researcher work, not a claim that they list a counted SOC 15-1221 inventory.

Build an AI Ethics Researcher resume on Resume Now

Write an AI Ethics Researcher resume, or one aimed at Natural Sciences Managers, instead of a blank template. Resume Now is a resume builder; we are not claiming a counted template set for this SOC.

Build an AI Ethics Researcher resume on Zety

An AI Ethics Researcher resume that names the actual tasks on this page, or the step-up title Natural Sciences Managers, beats a blank template when you apply.

What AI Ethics Researchers earn by state

This page does not show a state table, and the reason is worth stating: the Bureau of Labor Statistics does not publish a separate wage series for this job title, so there are no official state figures to show. Scaling the national median by a cost-of-living index would produce a number for every state, but it would be an estimate of living costs wearing a wage’s clothes, and PayCrunch would rather show you nothing than that.

What the national figures say: pay starts near $78,000, the median is $125,000, and the top of the range is $282,110. Those national figures are a PayCrunch estimate, not a Bureau of Labor Statistics published wage for this exact title.

If you want to see how far state pay can move for jobs the Bureau does publish state-by-state, the best-paying state for every occupation is a free open dataset, and the salary-by-state statistics page summarises the pattern across all 824 of them.

Free data. Use any of it.

PayCrunch publishes verified, BLS-sourced salary + AI-playbook data on 1,000+ professions — free, no signup.

Frequently asked
Will AI replace AI ethics researchers?
It is the opposite — the more AI ships, the more society needs people to evaluate it, and regulators are now mandating exactly this work. AI helps you run bigger evals and draft faster, but judging what counts as harmful, designing valid experiments, and standing behind independent findings are human responsibilities with real accountability. The risk is not replacement; it is being out-rigor'd by researchers who use AI to test at greater scale.
Isn't it a conflict to use AI to audit AI?
Only if you hide it or let the model grade itself. Disclose AI assistance, keep the model-under-test separate from any model you use to score, use held-out judges, and make everything reproducible. Used transparently, AI is just another instrument — used carelessly, it invalidates your results.
Do I need a PhD to reach the top of this field?
It helps for research-scientist roles at frontier labs, but the field also pays well for applied governance, red-teaming, and policy work where a strong public portfolio — reproducible evals, cited briefs, real audits — can matter as much as the degree. Build the portfolio either way.
How do I keep up when models change monthly?
Automate the boring part. Version your evals so re-running them on a new model is one command, use Perplexity/Elicit to track new literature weekly, and focus your human attention on new failure modes rather than re-reading everything. Reproducible tooling is how you stay current without drowning.
How does this work translate into pay at the top of the range?
The top of the band is frontier labs, big-tech responsible-AI teams, and leading think tanks — all hiring for a demonstrated ability to measure model behavior rigorously and shape governance. Reproducible evals, red-team results, and cited publications are the exact evidence those roles screen for.
Methodology & sources
  • Salary (median, 10th, top of the range) — U.S. Bureau of Labor Statistics, OEWS.
  • By state — the Bureau of Labor Statistics’ own state medians, limited to states employing at least 500 people in the occupation. No cost-of-living arithmetic is applied to a wage anywhere on this page.
  • The plays — PayCrunch's own step-by-step guidance using publicly available AI tools. Tool names/URLs are real and current as of August 2026; prompts written to work as-is. Verify any professional output before relying on it.

Sources