The research data scientist who checks their own work
$224,920top of the range in California · middle $120,230 / yr
AI is transforming this role
Data Scientist Researchs in the United States earn a median of $120,230 a year. Pay starts near $67,240. Pay reaches $224,920 at the top of the range in California, the best-paying state for this work among those with at least 500 people in the job.
Source: U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (Data Scientists, SOC 15-2051). Last checked 9 September 2026.
Entry level
$67,240
Top of the range · California
$224,920
Education
Master's or Doctoral degree
Wages — U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (Data Scientists). Top of the range is the highest state-level figure among states with at least 500 people in the job. AI-impact rating is PayCrunch's editorial assessment. Updated September 2026.
🆕 New & Trending AI Tools for Data Scientist ResearchReviewed September 2026
We track new AI-tool launches every week and refresh this list — here’s what’s gaining traction for Data Scientist Research work right now.
Claude CodeNEWFree / usage-based
Terminal coding agent that reads your repo, runs tests, and ships multi-file changes.
How a Data Scientist Research uses it: describe a feature and let it implement and test it across the codebase
OpenAI CodexNEWIncl. w/ ChatGPT plans
Agent that runs longer, deterministic multi-step coding jobs on its own.
How a Data Scientist Research uses it: delegate a well-defined build or migration and review the finished result
WindsurfNEWFree / $15 mo
Agentic IDE that keeps context across a whole project.
How a Data Scientist Research uses it: make large, coordinated changes without losing track of the codebase
AWS KiroNEWPreview / see site
Spec-driven coding agent that turns written specs into working code.
How a Data Scientist Research uses it: write the spec first and let it build to that spec
NotebookLMNEWFree / $7.99 mo
Google tool that answers questions grounded only in the documents you give it — with citations.
How a Data Scientist Research uses it: load your own manuals, policies, or PDFs and ask questions that stay accurate to the source
CursorFree / $20 mo
AI-native code editor that edits across an entire project.
How a Data Scientist Research uses it: describe a change in plain English and let it rewrite and refactor whole files
GitHub Copilot (Agent Mode)$10–19 mo
AI pair-programmer built into VS Code and GitHub that now completes multi-step tasks.
How a Data Scientist Research uses it: hand off a task and have it plan, edit multiple files, and open a pull request
ChatGPTFree / $20 mo
The most-used AI assistant — writing, analysis, research, and images from a plain-language chat.
How a Data Scientist Research uses it: draft emails and documents, summarize long files, and get instant answers to on-the-job questions
ClaudeFree / $20 mo
AI assistant known for careful writing, long-document analysis, and coding.
How a Data Scientist Research uses it: analyze big reports or spreadsheets and turn messy notes into clean, finished writing
A methods draft, still dense with comments, is pinned beside the whiteboard when the research data scientist opens the lab meeting. The group has results from the last experiment and a disagreement about whether the finding holds if one assumption is relaxed. Someone has to say which check comes next, which comparison would actually threaten the claim, and what the paper will be willing to conclude. That conversation, inside a university lab, a company research-and-development group, or a national laboratory, is the center of this seat.
The week orbits methods, evidence, and writing other researchers can inspect. A graduate degree is the common preparation, because the work asks you to design a study, defend it, and publish it. There is no occupational licence for the title. Hiring follows the papers and the seminar. Pay, when an offer arrives, is read from the figures in the salary section below.
Inside the lab meeting
Morning in a research group rarely starts with a dashboard for a sales target. It starts with a problem the study is built to answer, a dataset or a simulation that can speak to it, and a method that has to be right for that problem. The research data scientist reads what neighboring papers did, notices where their assumptions differ, and designs an experiment or an analysis that isolates the point under debate. Code is how the check is run. The notebook is a lab instrument, shared with the group, not a private scratch pad that disappears after one figure looks exciting.
Methods are the substance colleagues will attack first, in the friendly sense of a seminar. Why this model family, why this split of the data, why this baseline, what happens when a tuning choice changes, and which uncertainty belongs around the result. The scientist writes those choices down so a reader outside the room could rerun the argument. Negative results stay in the record. A pretty chart that hides a broken comparison is a setback, and the group is supposed to catch it before a draft leaves the building. Domain scientists, whether they study health, language, materials, or climate, bring the phenomenon. The research data scientist brings the method that can support a claim about it.
Papers are the durable product. A draft moves from an outline to a methods section that could stand alone, to results that match the code, to a discussion that does not overclaim. Internal review comes before any outside submission. Coauthors argue about a paragraph. The scientist who owns the analysis must be able to point at the exact run that produced a number in the text. In industry research groups the paper may be a technical report, a conference submission, or a methods note the next team will rely on. The audience is still other researchers. The success is a claim that survives scrutiny, not a button a customer clicks that afternoon.
The people around the bench shape the day. A principal investigator or a research director sets the program. Postdoctoral fellows and research engineers share code and critique. A lab manager or a program coordinator keeps data access and computing allocations in order. Seminars fill a standing slot on the calendar, and reading groups keep the group honest about new work. Some days are quiet implementation. Some days are a review where your favorite analysis is taken apart. Both belong to the job. The advice for someone entering is to seek a group that still does that taking-apart in the open.
Graduate study and other proof
A graduate degree is common for this seat. Departments of statistics, computer science, applied mathematics, and quantitative corners of the sciences and social sciences train people to design studies and to write them up. A master’s can be enough for some industrial research groups, especially paired with a strong thesis. A doctorate is the ordinary ticket for a university lab job that will lead projects and advise students, and it is frequent in corporate research labs as well. The degree is evidence of supervised research, not a licence from the state.
No government board issues a licence that authorizes someone to practice as a research data scientist. Employers look at the scholarship. A thesis, a preprint, a conference paper, or a careful technical report shows that you can carry a method from idea to defended result. Code and data, shared to the extent the project allows, let a hiring committee see whether the work is reproducible. Letters from advisors matter because research is apprenticeship-like: someone has watched you respond to criticism. Coursework alone, without a study you owned, is a thin packet for this particular seat.
Preparation during the degree is the research itself. Choose problems where the method is doing real work, keep a lab notebook your future self can read, and present at the group meeting before you feel ready. Learn to write the methods section early, while the choices are fresh. If you are already in industry and want this seat, look for an internal research-and-development group that publishes, and ask whether a joint methods project can become your bridge. A stack of predictive scores built only to move a weekly business metric will not, by itself, read as a research record. The record they want is the paper and the reasoning behind it.
What the group is buying
A research seat pays for methods, for experiments the group can defend, and for papers or technical reports other researchers can use. Bring that evidence to the interview. A model with no written method behind it describes a different kind of job.
How research groups hire
Watch university labs, corporate research organizations, national laboratories, and industrial groups attached to a larger company but measured on publications and prototypes of ideas rather than on this quarter’s operational report. Titles include research scientist, research data scientist, applied scientist inside a lab, and postdoctoral fellow. Read the announcement for papers, methods, and a group. If the text is only about shipping a score into a product decision, you are looking at a different seat even if the words data scientist appear.
The application is the scholarly record plus a letter that a human can stand to read. Name the problem, the method, and what the result changed in the group’s understanding. Link the paper and, when you may, the code. A job talk or a research seminar is the usual interview centerpiece: you present a study, and the room tries to break the method. Welcome that. Committees learn more from how you handle a flaw than from a seamless slide. Expect follow-up conversations with future collaborators about what you would study in their data, and about computing or data-access constraints you would need them to solve with you.
Probe the group’s real life one topic at a time so the conversation stays specific. Who decides the research program? Do people publish, and do they present outside the company? How is a methods disagreement settled? What happened to the last person in this seat? A lab that cannot describe its last paper, or a company group that has not written anything down in years, may be using a research title for ordinary production work. You can still take the job with open eyes. You should not be surprised later.
Timing often follows the academic calendar for campus jobs and a rolling calendar for industry labs. A finished thesis, a public talk, and a advisor who will speak plainly are worth more than a rushed application with none of those. If you need work before the doctorate is done, a research assistant role or an internship inside the kind of group you want is a legitimate step, provided you will touch methods and writing, not only data cleaning for someone else’s unnamed study. Keep your own notes. They become the next paper’s seed and the next letter’s detail.
A longer line of inquiry
The first research job is usually someone else’s program. You implement, you extend, you learn the group’s taste for evidence, and your name appears on papers you did not solely invent. That apprenticeship is the point. A later independent seat means you can propose the next study, recruit collaborators, and be trusted with the methods section of a flagship paper. In a university that path runs through postdoctoral work toward a research faculty or a staff scientist role. In a company lab it runs toward senior research scientist and, for some, a research director who sets themes for several people.
Depth beats a scattered publication list that never returns to the same problem. Groups remember the person who spent years making one method honest: better checks, clearer limits, a follow-up paper that answered the objection raised in the first review. Leadership, when it comes, is still intellectual. You critique kindly, you protect time for writing, and you decide which collaborations are real. Managing a large staff without remaining able to read a methods section is a different career. If you want to stay a research data scientist, keep a study you still understand line by line.
Moves between a campus and an industry lab are ordinary. What travels is the ability to pose a problem, choose a method, and write so strangers can test you. What you relearn is the mission: a company research group may need the idea to eventually matter to a product generation, while a university group may need it to matter to a field. Either way, the daily craft in this seat stays the experiment, the seminar, and the paper. Keep the code reproducible and the claims matched to the evidence. That is the promotion case, and it is also how the wider field decides whether to trust you the next time your name appears on a draft.
A research offer and the published range
A laboratory or a research-and-development group placing a salary on an offer to a data scientist in a research seat can read that offer against the May 2025 Occupational Employment and Wage Statistics series titled Data Scientists, SOC 15-2051. The national entry figure is $67,240, the national median is $120,230, and the high end of the published range in California is $224,920. California is where that high end is reported, among places the Bureau publishes. From $67,240 up to $120,230, the difference is $52,990. From the national median up to the California high end, the difference is $104,690.
Typical pay in a state is the state median, and it should not be confused with the high end. Washington’s median is $163,350, and that median stands $43,120 above the national median, the largest such step on this chart. California’s median is $141,590, a different number from the $224,920 high end in the same state. Maryland’s median is $136,370, New Jersey’s is $135,280, and Massachusetts shows $131,750. Louisiana’s median, $78,760, is the lowest published. A research hire comparing campuses or labs in these places can lay the offer on the matching median, and can mention Washington’s $43,120 gap when explaining why a Seattle-area lab’s typical pay sits higher than the national middle.
A new researcher still closely supervised, or a first industry-lab seat just after a master’s, may see an offer near $67,240. Once you have defended a study and other people rely on your methods section, the median of $120,230 is the fairer landmark, and $52,990 is the span you can talk through with the paper in hand. Say the numbers as annual pay. A postdoctoral-style title and a staff-scientist title can carry different offers inside the same lab. Match the figure to the independence, not to the prestige of the building. If the lab is in Maryland or Massachusetts, cite $136,370 or $131,750 as typical pay in that state, then ask where their letter sits relative to the national median as well.
Treat $224,920 as the high end of the published range in California, $104,690 above the national median. It is the right landmark for a scarce senior research post in that state, the sort of seat that leads a program and is hard to fill. It is a poor opening request for a first authorship under heavy supervision. In California, speak the median $141,590 when you mean typical pay, and speak $224,920 only when you mean the top of the published range. In New Jersey, the median to put on the table is $135,280. In Louisiana, $78,760 tells you typical pay is far nearer the national entry than the national median, so an offer there needs its own local reading.
Ask the group to state base pay for the year. Research roles sometimes add a summer salary, a stipend structure, or a company bonus. Count only dollars they will write down, then set the total beside $67,240, $120,230, and the California high end if that comparison is honest. The paper and the method are what the lab is buying. The salary is how you check that the letter matches the independence they described, using these published amounts and the state median that fits the city.
The top of Data Scientist Research pay — and how to get there with AI
$224,920what Data Scientist Research pay reaches in California
Highest state-level top-of-range annual wage for Data Scientists, among states with at least 500 people in the job. U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025.
And the role it leads to — Natural Sciences Managers — reaches $330,050 in California.
$67,240entry$120,230middle$224,920top end
Synthesising trend data into a recommendation is common; what marks the top of this range is the person who came back a quarter later, measured whether the recommendation worked, and published the answer whichever way it went.
Recommendations for action are produced constantly and almost never audited. Somebody synthesises business intelligence into a course of action, identifies an industry or geographic trend with strategy implications, generates the report for executives, and then the next study begins. The duties that would close that loop are on the list and routinely skipped: reviewing technical design documentation for accuracy, documenting specifications for the outputs, maintaining a library of model documents and reusable assets. Models now write the analysis narrative, which makes narrative cheap and verification valuable. A researcher who says out loud that something did not replicate ends up trusted with the questions that carry real money behind them.
Your playbook, by where you are now
Just startingWrite the prediction before the analysis
Record what you expect to find and what would show you were wrong, before you touch the data.
Pin every extract to a date and keep the query beside the finding, so a result can be reproduced instead of defended from memory.
Add a limitations paragraph to each report naming the population it does not cover.
Run heavy work on Apache Spark or Amazon Redshift against that pinned extract rather than whatever the table happened to hold that morning.
Ask Claude to argue against your conclusion using only your own figures, then answer the strongest objection inside the report.
What proves it: A finding somebody else reproduced from your recorded inputs.
Realistic span: the first year or two
A few years inBuild the follow-up into the study
Set a review date for every recommendation and schedule it in Apache Airflow so it cannot be quietly forgotten.
Design the measurement before the change happens, including a comparison group wherever one is possible.
Keep the library of model documents and templates current so methods stop being reinvented on each project.
Score how your trend calls aged: which held, which reversed, and what the reversals had in common.
Publish one internal replication of an earlier finding, including the parts that failed to hold.
What proves it: A scored record of your own past recommendations and how they turned out.
Realistic span: years three through six
ExperiencedSet the evidence bar
Define what your organisation requires before a finding may be called decision-grade, and apply it to your own output first and loudly.
Review other teams' technical design documentation for reporting solutions and say plainly when a design cannot support the claim being made on it.
Take the questions where no clean comparison exists, since that is where the seniority actually is.
Choose the modelling platform, Amazon Web Services AWS SageMaker or otherwise, on reproducibility grounds rather than on demonstrations.
California pays this work best, and managing scientific staff is the usual step across.
What proves it: An evidence standard adopted beyond your team, with rejected findings showing it has teeth.
Realistic span: seven years and beyond
The next 90 days
Find a recommendation your organisation acted on nine to twelve months ago and audit it. Get the data from before and after, work out what actually changed, and write two pages saying whether the expected effect appeared, whether something else explains it, and what you would do differently. Circulate it even if the answer is unflattering, particularly if it is. This is uncomfortable exactly once, and it changes how people treat your next report, because you become the researcher whose recommendations come with a record rather than a forecast. It also teaches you which of your habits produce findings that survive contact with reality, which is knowledge no amount of new analysis will give you.
Wage figures: BLS OEWS, May 2025. The playbook is PayCrunch editorial guidance, not a guarantee of pay or placement.
Every figure is the national median from the U.S. Bureau of Labor Statistics (OEWS) shown on that role’s own page.
Never used AI before? Start here (2 minutes).
Put an AI coding copilot inside your daily workflow first. Add GitHub Copilot to your IDE or move to Cursor, and use ChatGPT Advanced Data Analysis or Julius AI for fast, conversational exploration of a dataset. The copilot writes the boilerplate so you spend your hours on the framing and the modeling that only you can do.
For learning and free reference, work through Kaggle competitions and datasets, pull pretrained models from Hugging Face, and run everything in a free Google Colab notebook with Gemini as your in-notebook helper. Keep confidential data out of all of these; learn on public and synthetic data.
The one rule, forever: Never paste proprietary data, PII, or unpublished results into a consumer AI tool; use enterprise zero-retention endpoints or local models for anything confidential. Treat AI-written code and statistics as a draft: re-run it, check the assumptions, and confirm every citation, because hallucinated references and silently wrong stats are the fastest way to burn your credibility.
The plays — exact steps, exact prompts
Do these in order. Each one is copy-paste ready. You do not need to know anything about AI going in.
1
Pair-program every pipeline instead of typing it
Why this pays: Your output is measured in shipped analyses and models. An AI coding copilot cuts the time from question to working notebook roughly in half, so you run more experiments per quarter, the visible productivity that gets you to staff and principal band.
CursorGitHub CopilotClaude Code
1
Move day-to-day work into Cursor (or add GitHub Copilot to VS Code or JupyterLab). Let it write the boilerplate, data loading, joins, and plotting scaffolds, while you own the logic and the choices.
2
When you hit a gnarly transformation or a slow query, describe it in plain English and let the copilot draft it, then review.
Copy-paste this prompt
I have a pandas DataFrame [df] with columns [list columns + dtypes]. I need to [describe the transformation, e.g. compute a 7-day rolling median per user_id, handling gaps]. Write vectorized pandas (no explicit Python loops), explain the approach in 3 bullets, and flag the edge cases (NaNs, unsorted timestamps, duplicate keys) I should test.
Always re-run the generated code on a small known sample and check the edge cases it names before trusting it at scale.
3
Use Claude Code in the terminal to refactor a messy research script into a reproducible pipeline with functions, tests, and a config file.
What you'll haveTwo to three times more experiments shipped per quarter with cleaner code, the throughput that justifies a senior or staff title.
2
Turn a week of literature review into an afternoon
Why this pays: Framing the right question on top of what is already known is what makes research trusted. Doing it in hours instead of weeks means you start more projects and avoid dead ends, leverage that compounds into principal-level impact.
ElicitConsensusPerplexity
1
Start a question in Elicit to pull the relevant papers into a structured table (method, sample size, effect, limitations) instead of reading 40 PDFs by hand.
2
Sanity-check the consensus and the disagreements before you commit to a modeling approach.
Copy-paste this prompt
Summarize the state of evidence on [method X vs method Y for problem Z]. Give me: (1) where the literature agrees, (2) the main open disagreements, (3) which approach is standard for [my data size / domain], and (4) the 3 papers I must read. Note anything that would change the answer for [my specific constraint].
Use Consensus or Elicit for the citation-backed version, and verify each paper actually says what the summary claims before you rely on it.
3
Keep a running Perplexity thread for fast, cited lookups of unfamiliar methods or libraries mid-project.
What you'll haveA defensible method choice grounded in the literature, reached in an afternoon: fewer dead-end projects, more credible results.
3
Explore and model at the speed of thought
Why this pays: The exploratory phase is where most hours vanish. AI data-analysis copilots let you interrogate a dataset conversationally, so you reach the signal faster and can take on more, higher-stakes questions.
Julius AIHex MagicChatGPT Advanced Data Analysis
1
Drop a dataset into Julius AI or ChatGPT Advanced Data Analysis and ask for the first-pass EDA, distributions, missingness, correlations, and obvious leakage, in one prompt.
2
In Hex Magic, generate the SQL or Python cell from a plain-English ask, then keep the notebook as the reproducible artifact.
Copy-paste this prompt
Here is the schema of [dataset]. Do a rigorous first-pass EDA: flag missingness and outliers per column, show the target's distribution, list the top candidate features with a one-line rationale each, and warn me about any target leakage or train/test contamination risks given [describe how the data was collected].
Leakage warnings are only a prompt to investigate; confirm the data-generating process yourself, because the model cannot see how the table was actually built.
What you'll haveDays of exploratory grunt work compressed into hours, freeing you for the modeling and framing that actually move a project.
4
Own reproducibility and experiment tracking
Why this pays: The data scientist leadership trusts is the one whose results reproduce. Owning the experiment-tracking and reproducibility stack makes you the technical anchor of the team, the profile that gets promoted to staff and principal.
Weights & BiasesDVCMLflow
1
Instrument every model run with Weights & Biases (or MLflow) so metrics, configs, and artifacts are logged automatically. No more 'which version got that number?'
2
Version data and pipelines with DVC so any result can be regenerated from scratch.
3
Have AI draft the reproducibility scaffolding so the whole team adopts it.
Copy-paste this prompt
Design a lightweight experiment-tracking and reproducibility setup for a small research data-science team using [W&B / MLflow] and DVC. Include: what to log per run, a repo structure, a config-driven training entrypoint, and a checklist that guarantees a result from 6 months ago can be reproduced. Keep it low-friction so people actually use it.
Adapt to your stack and security rules; never let a tracking tool upload confidential data to a shared cloud project without approval.
What you'll haveResults that reproduce on demand and a team that runs on your system, the reliability that earns the senior technical seat.
5
Make results impossible to ignore
Why this pays: Analyses that are not understood do not get you promoted. Using AI to turn findings into crisp, decision-grade narratives and visuals is how your work reaches, and persuades, the people who set comp.
ClaudeQuartoChatGPT
1
Draft the executive narrative of a result with Claude or ChatGPT: you supply the numbers and caveats, it structures the story for a non-technical decision-maker.
2
Turn the analysis into a polished, reproducible report or dashboard with Quarto.
Copy-paste this prompt
I ran [analysis] and found [key results + the 2 most important caveats]. Write a one-page brief for [VP of Product / non-technical execs]: the decision it informs, what we found in plain language, the confidence level and caveats stated honestly, and a clear recommendation. No jargon, no overclaiming.
Never let the AI inflate certainty; you are accountable for every claim, so keep the caveats it tends to smooth over.
What you'll haveFindings that land with decision-makers and get attributed to you, the visibility that drives promotions and raises.
6
Ship models with modern ML copilots
Why this pays: Moving from analysis to models that run in production is the jump to the highest-paid data-science work. AI copilots let a single scientist prototype, tune, and ship models that used to need a team.
Hugging Face AutoTrainPyTorchWeights & Biases
1
Prototype a baseline fast with Hugging Face models and AutoTrain before hand-rolling anything in PyTorch.
2
Use an AI copilot to design the training loop and debug convergence, tracked in W&B.
Copy-paste this prompt
I'm training [model type] on [data + task]. Draft a PyTorch training loop with proper train/val/test splits, early stopping, and Weights & Biases logging. Then give me a checklist to debug [describe symptom, e.g. validation loss diverging] and the 3 most likely causes for this kind of model.
Validate the split and metrics yourself; a copilot will happily write a loop that quietly leaks the test set.
What you'll haveProduction-grade models shipped solo, the skill set behind the top-of-range, staff-level data-science salary.
Your 12-month sequence to the top of the range
How the plays above stack into a path from median pay toward the $224,920 tier.
Month 1
Put a coding copilot (Cursor or Copilot) in your daily workflow and route literature reviews through Elicit and Consensus. Measure the hours saved.
Months 2-3
Move EDA and modeling into AI data-analysis copilots, and stand up W&B and DVC so every experiment is tracked and reproducible.
Months 3-6
Ship a model end to end with Hugging Face and PyTorch, and start writing AI-assisted decision briefs so leadership sees the work.
Months 6-12
Own the team's reproducibility and experiment-tracking standard, and take the lead on one high-stakes research question, the staff/principal move.
Gear for this job
As an Amazon Associate, PayCrunch earns from qualifying purchases. Links to books and tools are for the job on this page; we only recommend what we’d use in the work.
Same live O’Reilly 3rd already on data-scientist / python-developer / market-research-analyst / biochemist / botanist. This page’s pair-program play prompt is I have a pandas DataFrame / Write vectorized pandas. Not CompTIA Data+ and not leftover Ross Exam P (that is actuary).
Next steps for a Data Scientist Research
Some links below are affiliate or partner links. PayCrunch may earn a commission if you enroll or subscribe through them, at no extra cost to you. Wage figures on this page still come from the Bureau of Labor Statistics, not from these programs.
Data Scientist Research work is specific enough that a stamped 'check out these courses' block would be noise. BLS files this work as Data Scientists (SOC 15-2051). O*NET Job Zone 4 is typical: a bachelor's degree, so the honest next credential is a professional certificate or bachelor's-level coursework — not a random catalog dump.
Data Scientist Researchs in this dataset list AJAX among the tools in use, so a program that names that stack is a better fit than a survey course.
The next title this dataset points at is Natural Sciences Managers; a credential aimed that way is a clearer step than another year in the same seat.
Coursera search for data science — a professional certificate or bachelor's-level coursework that lines up with computing, not a generic professional-development aisle.
FlexJobs screens remote, hybrid, freelance, and flexible listings so you are not wading through unverified ads. This is a job-board search for Data Scientist Research work, not a claim that they list a counted SOC 15-2051 inventory.
Write a Data Scientist Research resume, or one aimed at Natural Sciences Managers, instead of a blank template. Resume Now is a resume builder; we are not claiming a counted template set for this SOC.
A Data Scientist Research resume that names the actual tasks on this page, or the step-up title Natural Sciences Managers, beats a blank template when you apply.
What Data Scientist Researches earn by state
These are the Bureau of Labor Statistics’ own figures for Data Scientists, state by state — not a cost-of-living adjustment applied to the national number. Only states employing at least 500 people in the occupation are shown, because a state median drawn from a handful of workers is noise rather than a signal.
Washington
$163,350
highest of them · +36% vs the national median
Louisiana
$78,760
lowest of the 40 states and D.C. that qualify · -34% vs the national median
The same job pays $84,590 more a year at the median in Washington than in Louisiana — 107% higher. That gap is what the Bureau measured, before any question of what it costs to live in either place. The top-of-range figure quoted at the head of this page, $224,920, is a different statistic in a different place: it is the 90th-percentile wage in California. The state that pays the typical worker most and the state where the best-paid go highest are not always the same one.
Source: U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025, SOC 15-2051. 40 states and D.C. clear the 500-employee reporting floor for this occupation; those below it are left out rather than shown with a wide error band.
Free data. Use any of it.
PayCrunch publishes verified, BLS-sourced salary + AI-playbook data on 1,000+ professions — free, no signup.
No, but it changes the job. Framing the question, causal reasoning, domain judgment, and accountability for the result stay human; routine coding and first-pass EDA get automated. The data scientists who use AI run three times the experiments; the ones who ignore it end up competing on speed they cannot win.
Is it safe to paste my data into ChatGPT or Claude?
Not proprietary data, PII, or unpublished results. Use enterprise zero-retention tiers or a local model for anything confidential. Public, synthetic, or already-published data is fine, and that is plenty for learning and prototyping.
Won't AI-written code introduce subtle bugs and data leakage?
Yes, and that is the real risk. Copilots write code that runs but is silently wrong: leaked test sets, mis-specified splits, broken group boundaries. Treat every generated block as a draft, re-run it on a known case, and check the assumptions before you trust a number.
Do I still need to learn statistics and ML deeply if AI can do it?
More than ever. You are the one who catches when the AI picked the wrong method, ignored an assumption, or fabricated a citation. Depth is exactly what makes you able to direct the tools instead of being misled by them.
How does this actually raise my salary?
Three levers: throughput (more shipped experiments and models), reliability (reproducible results leadership trusts), and communication (findings that reach decision-makers). Together they define the staff/principal band and make you competitive for the higher-paying industry research roles at the top of it.
Methodology & sources
Salary (median, 10th, top of the range) — U.S. Bureau of Labor Statistics, OEWS.
By state — the Bureau of Labor Statistics’ own state medians, limited to states employing at least 500 people in the occupation. No cost-of-living arithmetic is applied to a wage anywhere on this page.
The plays — PayCrunch's own step-by-step guidance using publicly available AI tools. Tool names/URLs are real and current as of August 2026; prompts written to work as-is. Verify any professional output before relying on it.