The ETL developer whose framework the team runs on
$178,580top of the range in New York · middle $104,620 / yr
High AI exposure
ETL Developers in the United States earn a median of $104,620 a year. Pay starts near $60,230. Pay reaches $178,580 at the top of the range in New York, the best-paying state for this work among those with at least 500 people in the job.
Source: U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (Database Administrators, SOC 15-1242). Last checked 9 September 2026.
Entry level
$60,230
Top of the range · New York
$178,580
Education
Bachelor's degree in Computer Science
Wages — U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025 (Database Administrators). Top of the range is the highest state-level figure among states with at least 500 people in the job. AI-impact rating is PayCrunch's editorial assessment. Updated September 2026.
🆕 New & Trending AI Tools for ETL DeveloperReviewed September 2026
We track new AI-tool launches every week and refresh this list — here’s what’s gaining traction for ETL Developer work right now.
Claude CodeNEWFree / usage-based
Terminal coding agent that reads your repo, runs tests, and ships multi-file changes.
How an ETL Developer uses it: describe a feature and let it implement and test it across the codebase
OpenAI CodexNEWIncl. w/ ChatGPT plans
Agent that runs longer, deterministic multi-step coding jobs on its own.
How an ETL Developer uses it: delegate a well-defined build or migration and review the finished result
WindsurfNEWFree / $15 mo
Agentic IDE that keeps context across a whole project.
How an ETL Developer uses it: make large, coordinated changes without losing track of the codebase
AWS KiroNEWPreview / see site
Spec-driven coding agent that turns written specs into working code.
How an ETL Developer uses it: write the spec first and let it build to that spec
NotebookLMNEWFree / $7.99 mo
Google tool that answers questions grounded only in the documents you give it — with citations.
How an ETL Developer uses it: load your own manuals, policies, or PDFs and ask questions that stay accurate to the source
CursorFree / $20 mo
AI-native code editor that edits across an entire project.
How an ETL Developer uses it: describe a change in plain English and let it rewrite and refactor whole files
GitHub Copilot (Agent Mode)$10–19 mo
AI pair-programmer built into VS Code and GitHub that now completes multi-step tasks.
How an ETL Developer uses it: hand off a task and have it plan, edit multiple files, and open a pull request
ChatGPTFree / $20 mo
The most-used AI assistant — writing, analysis, research, and images from a plain-language chat.
How an ETL Developer uses it: draft emails and documents, summarize long files, and get instant answers to on-the-job questions
ClaudeFree / $20 mo
AI assistant known for careful writing, long-document analysis, and coding.
How an ETL Developer uses it: analyze big reports or spreadsheets and turn messy notes into clean, finished writing
The warehouse is late and someone is already asking
An analyst opens the morning dashboard and the numbers stop at yesterday. A finance partner messages the channel. A source system changed a field name over the weekend and nobody sent a note. The ETL developer is the person who moves data from those source systems into the warehouse, on a schedule the business has learned to trust. When the load breaks, the job is to see which step failed, repair the pipeline, rerun what can be rerun, and tell the people waiting what is safe to use. The title varies. Pipeline developer, analytics engineer, data engineer on a warehouse team. The work is the movement: extract, reshape, load, and the schedule that is supposed to make that movement boring.
Source systems are whatever the company already runs. An orders database. A payroll export. A spreadsheet a regional manager still updates. A vendor file that arrives when it arrives. The warehouse is where those feeds are supposed to agree, so a report does not depend on someone pasting columns together. You learn the meaning of the fields, not only their names. A column called status can mean three different things in three systems. Your pipeline has to pick one meaning and keep it. If you guess, the dashboard will be confident and wrong, which is worse than being late.
Pipelines, schedules, and a load you can explain
A pipeline is the path from a source to a table people query. Extract pulls the records. Transform cleans types, joins keys, and applies the business rules the warehouse agreed to. Load writes the result where analysts expect it. Some paths run every night. Some run through the day. Some start when a file lands. The schedule is part of the product. A pipeline that works only when you remember to click it is a hobby. You document when it runs, what it depends on, and what "done" looks like so another person can tell success from a silent skip.
Broken loads have a small set of shapes you learn to recognize. The source was empty. The source was duplicated. A key that used to be unique now repeats. A date arrived in a new format. A job upstream finished late, so yours started on stale input and stamped the warehouse as fresh. Your first task is to stop the bleeding: mark the data as incomplete if people are about to make a decision on it. Then you find the step. Then you fix the mapping or the dependency, rerun, and check row counts against the source, not against your hope. A green job icon proves little by itself. A count that matches, and a spot check of the awkward records, is closer to proof.
The rest of the week is prevention. You add a check that fails loudly when a file is missing. You agree with the team that owns the source app on how a column change will be announced. You write the pipeline so a rerun does not double the orders. You keep a short note on each feed: where it comes from, who to call, and which reports will lie if it stalls. Analysts should not have to discover a broken load by accident in a meeting. If they do, the schedule and the alert were part of the defect, not only the transform.
You sit between people who own operational systems and people who read the warehouse. The app team cares that checkout works. The finance team cares that yesterday's revenue is in the table before the standup. You translate. When a source cannot deliver a field you were promised, you say so early, and you adjust the pipeline instead of hiding a null behind a default that flatters the chart. Good ETL work is visible when it fails and quiet when it works. The quiet days are the ones you designed.
There is no licence, so the proof is the pipeline
No state licenses ETL developers. There is no board that grants a card and no national exam you must hold before you schedule a load. Employers treat the skill as professional practice without an occupational licence. What they accept instead is evidence you have moved real data and noticed when it went wrong. A bachelor's degree in computer science, information systems, or a related field is common. Plenty of people arrive from an analyst seat, a software role, or a business team that got tired of repairing the same spreadsheet. The degree helps. A pipeline you can walk through helps more.
Preparation is SQL you can use without a cheat sheet for joins and filters, plus one orchestration or pipeline tool you have actually scheduled. Learn how a warehouse table is modeled well enough to land data where analysts can query it without a rescue project. Learn to read a failure log. Practice on a public dataset or a carefully fake one: pull from two sources, conform a key, load on a timer, and break the source on purpose so you can show how you noticed. Vendor certificates exist for cloud platforms and database products. They can support a resume when the posting names that platform. They do not replace a story about a load you repaired. If a certificate's outline is mostly product trivia, spend the time on a pipeline instead.
What stands in for a licence
SQL, a scheduled pipeline, and a clear account of a broken load. No state board issues an ETL licence. A cloud certificate is optional proof, useful when the stack matches the job.
How teams hire the person who owns the feed
The roles sit on data teams, inside analytics groups, and sometimes in IT at a company large enough to keep a warehouse. Titles drift. Read the task list. If it says pipelines, schedules, source systems, and failed jobs, it is this work even when the posting says data engineer. If it is mostly building the application that creates the data, it is a different seat. Apply where you can talk about movement into a warehouse. Ignore adjectives like "ninja" and look for who gets paged when the morning table is empty.
Hiring managers screen for a walkthrough. They want the sources you have touched, the schedule you owned, and one failure you can narrate without blaming a vendor for the entire hour. Be specific. Name the kind of break, what you checked first, how you confirmed the rerun, and what you changed so the same break would announce itself. Bring a short diagram if you have one. A portfolio repository helps when the data is fake or public. Do not paste an employer's schema into a public repo to look impressive. Discretion about live customer data is part of the job, and interviewers notice if you treat it casually.
Expect a practical conversation more than a trivia contest. You may be asked to sketch how you would land a daily orders file, what you would do if the file arrived twice, and how you would tell an analyst to hold off on the number. Think out loud about keys, late data, and alerts. Ask them how many pipelines are already in production, who changes a source system, and whether failed loads wake someone up. A team with no alert and no owner is a team that will make the broken load your personality. That can be a good first job if they will teach you. It is a bad job if they want a hero and a silent warehouse.
From running someone else's job to owning a domain
Juniors usually inherit pipelines. You learn the scheduler, you take the small breaks, and you resist rewriting everything in your favorite tool during week two. The win in that season is reliability: jobs that finish, checks that fire, and notes another person can follow when you are out. You sit with an analyst once a week and ask which table they stopped trusting. Trust repaired on a single feed is a better early story than a new framework nobody asked for.
The next step is a domain. You own the feeds for finance, or for the product events, or for the vendor files that feed inventory. You know the source owners by name. You design the new pipeline when a system is replaced, including the period when old and new overlap and the numbers must be compared. You set conventions the team can share: how dates are stored, how a rerun behaves, how a breaking change is announced. People start to bring source problems to you because your answer includes the downstream reports, not only the job that failed.
Leads stop being the fastest fixer and start being the person who decides what "good" looks like for the platform. You review other people's pipelines. You argue for alerts that match real decisions, not for noise. You plan the cutover when a source system retires. Some developers later move toward modeling the warehouse more broadly, or toward the product side of analytics. The ETL path itself stays valuable as long as companies keep more than one system of record. Keep a private log of feeds you have owned and outages you have closed. That log is your promotion packet, and it is more persuasive than a list of tools you have installed.
Where these dollars come from
The figures are Occupational Employment and Wage Statistics for May 2025. They are published under Database Administrators, a series wider than the developer who moves data from source systems into a warehouse on a schedule. Entry pay is $60,230. The national median is $104,620. At the far end, the published figure is $178,580 in New York, in the locations the Bureau could include when it reported a high end. A state median is a different statistic. Utah's median is $135,750, which is $31,130 above the national median.
From entry to the national median is $44,390. From the national median to the New York high end is $73,960. Massachusetts posts a median of $129,300. New Jersey is $125,860. Maryland is $124,300. The District of Columbia is $118,540. Kentucky's median, $87,390, is the low end of the published state medians here. The gap from Utah's median to Kentucky's is $48,360. Compare medians when you are comparing places. Compare $178,580 only when you mean the high end of the published range in New York. Treat that New York figure as the high end of the published range. A state median is a separate comparison, and these figures do not include a New York median to put beside a typical offer there.
Using the gap that matches the job
Hold a junior offer, still running pipelines someone else designed, next to $60,230. Hold an offer for someone who owns a domain of feeds, the schedule, and the broken-load response next to $104,620. The $44,390 between those two national figures is the size of that step in the data. If the company wants on-call for failed loads and design of new pipelines, and the offer is still sitting on the entry number, say so with the duties named. Do not wave $178,580 at a first warehouse job. That $73,960 above the national median describes the far published edge in New York.
Geography needs its own sentence. Utah's median of $135,750 is the high state median in this set, $31,130 above the national median. Massachusetts at $129,300, New Jersey at $125,860, Maryland at $124,300, and the District of Columbia at $118,540 are the other medians worth putting beside an offer in those places. Kentucky at $87,390 is the published low median. The $48,360 between Utah and Kentucky is a location gap, not a skill gap you can claim by rewriting a resume. If you are moving, compare the offer with the median of the state you are entering, then decide whether rent changes the meaning of that number. The Bureau figures do not adjust for housing.
Keep bonus, on-call pay, and equity off the base until you have compared the base with the right figure. A title bump with no change in who gets called when a load fails does not justify a jump from the median toward the New York high end. A real expansion is a new domain, review of other developers' pipelines, or responsibility for the schedule the company bets a meeting on. Use $60,230, $104,620, or the state median that matches the offer. Use $178,580 only with the label high end of the published range, and only when the role and the place make that edge a serious comparison. The work you are pricing is still the same: sources, a warehouse, a schedule, and a load that broke.
The top of ETL Developer pay — and how to get there with AI
$178,580what ETL Developer pay reaches in New York
Highest state-level top-of-range annual wage for Database Administrators, among states with at least 500 people in the job. U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025.
And the role it leads to — Software Developers — reaches $272,670 in California.
$60,230entry$104,620middle$178,580top end
Shipping another pipeline keeps an ETL developer fully booked; building the thing every other pipeline is written against is what puts one at the top of this range.
Most teams accumulate feeds the way a garage accumulates tools: each one written by whoever needed it, each with its own retry logic, its own idea of a bad row, its own untested assumption about the source schema. Writing the logical and physical database descriptions, testing changes before they land, specifying user access per segment, and answering the same questions from junior staff are all listed parts of this job, and all of them scale badly one feed at a time. Assistants that write code have made the accumulation worse, not better, since a pipeline can now be generated in an hour by someone who cannot debug it at three in the morning. The developer who supplies the framework and the tests decides what that generated code has to satisfy.
Your playbook, by where you are now
Just startingMake one pipeline behave properly
Take your least reliable feed and make it restartable, so a failure halfway through can be rerun without duplicating rows.
Add contract checks on the source: expected columns, types and row counts, failing loudly rather than loading rubbish into Amazon Redshift.
Write the physical and logical descriptions of the target tables down, including why each key was chosen.
Get every job definition into version control, and make deployment a script rather than a person with credentials.
Use Cursor or GitHub Copilot for the boilerplate transformation code, then write the test that proves the output before you trust a line of it.
What proves it: One feed that can fail, be rerun and land identical data, with tests that catch a schema change.
Realistic span: the first two years
A few years inGeneralise it into something others use
Turn your restart, logging and validation patterns into a shared library every new job is built on.
Build the backfill command, because reprocessing history by hand is where weekends and data quality both go.
Generate lineage from the job definitions themselves, so which report breaks when a column changes stops being folklore.
Make adding a straightforward source a configuration change an analyst can make, with your framework enforcing access levels and validation.
Publish the standard for what a job must emit before it runs in production, and enforce it in review rather than in conversation.
What proves it: New feeds in your team built on your framework by people who did not ask you first.
Realistic span: years three through six
ExperiencedOwn the platform and the people on it
Set the platform direction across streaming and batch, whether that means Apache Kafka, Ab Initio or Informatica Corporation PowerCenter, and write down the reasoning.
Design the access model per data segment and the retention rules, since data nobody governs eventually becomes somebody's incident.
Support and train the junior engineers deliberately, with office hours and reviews, so the framework outlives your attention.
Move toward software engineering roles, the natural step up from this work, and note that New York prices this occupation highest.
What proves it: A platform standard other teams adopt and engineers who were trained on it, not just handed it.
Realistic span: year seven onward
The next 90 days
In the next ninety days, pick the feed that pages someone most often and make it the prototype for everything that follows. Give it a real contract with its source: which columns must exist, what types they hold, what row count range is plausible, and what happens when any of that is violated. Make it idempotent, so a rerun produces the same table rather than doubled rows. Put the job definition in version control and write two tests that fail when the source changes shape. Then extract whatever you built into something a colleague can import for their own job, and watch them use it. An ETL developer with a framework other people build on is not judged by how many pipelines they wrote that quarter.
Wage figures: BLS OEWS, May 2025. The playbook is PayCrunch editorial guidance, not a guarantee of pay or placement.
Every figure is the national median from the U.S. Bureau of Labor Statistics (OEWS) shown on that role’s own page.
Never used AI before? Start here (2 minutes).
Open your transformation layer's own AI first. If you use dbt, turn on dbt Copilot to generate models, tests, and docs from your schema; if you're on Databricks or Snowflake, use the Databricks Assistant or Snowflake Copilot to write and explain SQL in context. Then verify every generated query against sample data and row counts before it touches a real table.
For everything else - refactoring, debugging, learning a new engine - use GitHub Copilot or Cursor in your IDE and Claude for reasoning through a gnarly transformation. Practice on synthetic data or public datasets; never paste production PII, keys, or connection strings into a public tool.
The one rule, forever: AI-written SQL and pipeline code is a draft that can silently corrupt data - wrong joins fan out rows, bad type casts drop precision, and a mis-scoped incremental can double-count revenue. Never merge machine-generated transforms without reviewing the logic, running them against sample data, and checking row counts and totals. Never paste real customer PII, credentials, or production connection strings into a consumer AI tool; use synthetic data and schemas only, and keep governed data in approved systems.
The plays — exact steps, exact prompts
Do these in order. Each one is copy-paste ready. You do not need to know anything about AI going in.
1
Ship pipelines faster with AI-generated dbt models and tests
Why this pays: The routine work of writing transforms and tests is exactly what AI does best - so the way to stay valuable is to do it in a fraction of the time and spend the surplus on design and quality. Shipping correct, well-tested models faster is what earns the senior 'analytics engineer' title and its pay.
dbt CopilotGitHub CopilotClaude
1
Use dbt Copilot (or GitHub Copilot in your dbt project) to generate staging and mart models from source schemas, then refactor the SQL yourself for correctness and performance - the AI drafts, you own the logic.
2
Auto-generate the tests you'd normally skip under deadline.
Copy-paste this prompt
You are a senior analytics engineer. For this dbt model [paste model SQL with synthetic column names], write dbt tests in schema.yml covering: primary-key uniqueness and not-null, accepted values for [status column], relationships to [dim table], and a reasonable freshness check. Explain what each test protects against and where the model could silently produce wrong numbers.
Use synthetic schemas only. Run every generated test against sample data and confirm row counts - an untested AI transform can quietly corrupt a metric.
3
Have AI draft the model documentation and column descriptions so your project stays self-documenting without eating your afternoon.
What you'll haveCorrect, well-tested, documented pipelines shipped in a fraction of the time - the productivity that promotes you from writing transforms to designing the platform.
2
Modernize legacy ETL - migrate Informatica/SSIS to a modern stack
Why this pays: Migrations off legacy tools (Informatica, SSIS, hand-rolled stored procs) onto dbt, Snowflake, or Databricks are high-budget, high-visibility projects. AI that reads legacy logic and translates it slashes the timeline - and being the person who leads a migration is a direct route to the top of the band.
ClaudeGitHub CopilotSQLMesh
1
Feed legacy mapping logic (sanitized) to Claude to reverse-engineer what a pipeline actually does before you rebuild it - AI is fastest at explaining code no one remembers writing.
2
Translate legacy transforms into modern, testable SQL.
Copy-paste this prompt
Here is a legacy [SSIS/Informatica] transformation described in pseudocode with synthetic column names: [paste sanitized logic]. Convert it into a clean dbt model (or SQLMesh model) for [Snowflake]. Preserve the exact business logic, flag any implicit type conversions or null-handling that could change results, and list the reconciliation checks I should run to prove old and new outputs match.
Migrations live or die on reconciliation. Run row-count and value-level diffs against the legacy output before cutover; never assume the AI preserved edge cases.
3
Build a parallel-run and reconciliation harness so old and new pipelines run side by side until the numbers tie out exactly, then cut over.
What you'll haveA faster, lower-risk migration you can lead end to end - the kind of flagship project that lifts an ETL developer into senior data-engineer comp.
3
Guarantee data quality with AI-generated tests and anomaly detection
Why this pays: Trust is the product. The moment a dashboard shows a wrong number, the business stops trusting the data team. Engineers who make pipelines provably reliable - with tests, contracts, and anomaly detection - become indispensable, and indispensable is what pay at the top of the range rewards.
Great Expectationsdbt (data contracts)Databricks Assistant
1
Use AI to expand test coverage across the warehouse - Great Expectations suites and dbt contracts - so schema changes and bad loads fail loudly in CI instead of silently in a report.
2
Have AI design anomaly checks for your real metrics.
Copy-paste this prompt
Act as a data reliability engineer. For a daily [orders] fact table, propose data-quality monitors that would catch: sudden row-count drops or spikes, a source that stopped loading, duplicate keys, revenue totals outside expected range, and late-arriving data. Give me the check logic (SQL or Great Expectations) and sensible thresholds, and explain the false-positive risk of each. Synthetic schema only.
Tune thresholds against real historical volume before enabling alerts - a noisy monitor gets muted, and a muted monitor protects nothing.
3
Wire the checks into CI and orchestration so a failing test blocks the merge or halts the load - quality enforced by pipeline, not by hope.
What you'll havePipelines the business provably trusts - the reliability reputation that makes you the engineer they can't afford to lose.
4
Cut warehouse cost and query time with AI optimization
Why this pays: Cloud data bills are enormous and rising, and the engineer who visibly cuts them by tuning queries, clustering, and warehouse sizing delivers dollars straight to the P&L. That visible savings is one of the easiest ways to justify a raise and a platform-owner role.
Snowflake CortexdbtClaude
1
Use Snowflake Copilot/Cortex or the Databricks Assistant to explain slow query plans and suggest rewrites, clustering keys, and incremental strategies - then validate the savings on real workloads.
2
Have AI rewrite an expensive transform for cost and speed.
Copy-paste this prompt
Here is a slow, expensive warehouse query and its plan [paste sanitized SQL + plan]. Rewrite it to reduce scanned data and compute cost on [Snowflake/BigQuery]: suggest incremental materialization, pruning/clustering, and any redundant joins or full scans to remove. Explain the tradeoffs and how to confirm the results are identical to the original.
Prove correctness first, then measure cost. Always confirm the optimized query returns identical results before shipping a rewrite.
3
Convert your heaviest full-refresh models to incremental and set up a cost dashboard so savings are visible to whoever signs your review.
What you'll haveA measurable drop in the data bill and faster pipelines - a dollar-denominated win that makes the raise conversation easy.
5
Own the platform - orchestration, lineage, and the semantic layer
Why this pays: The top of the band belongs to whoever owns the data platform, not the individual pipelines. AI lets you take on orchestration, lineage, CI/CD, and a governed semantic layer without a bigger team - and platform ownership is exactly the scope that pays $179k and above.
Apache Airflowdbt (semantic layer)GitHub Copilot
1
Use Copilot to write and maintain Airflow DAGs and CI/CD so pipelines deploy, test, and schedule themselves - orchestration is what turns a pile of scripts into a platform.
2
Stand up a governed dbt semantic layer so metrics are defined once and consumed everywhere, ending the 'why do two dashboards disagree?' fire drills.
3
Have AI help you design the architecture, not just the code.
Copy-paste this prompt
Act as a principal data engineer. Review my proposed pipeline architecture: [describe sources, ingestion, transformation, orchestration, and consumers at a high level - no secrets]. Critique it for reliability, cost, data-quality gaps, and single points of failure, and suggest where to add contracts, lineage, and backfill safety. Give me the three highest-impact improvements.
Use AI to pressure-test design decisions, then validate against your real SLAs and volumes - architecture tradeoffs are context-specific.
What you'll haveOwnership of the whole data platform instead of a queue of tickets - the scope and title that command pay at the top of the range.
Your 12-month sequence to the top of the range
How the plays above stack into a path from median pay toward the $178,580 tier.
Month 1
Turn on dbt Copilot (or Snowflake/Databricks Assistant) and GitHub Copilot. Rebuild one existing model with AI, add full tests and docs, and verify row counts and totals match.
Months 2-3
Expand AI-generated data-quality tests and anomaly monitors across your most-used tables and wire them into CI so bad data fails loudly.
Months 3-6
Lead an optimization or migration win: cut warehouse cost with AI query tuning, or translate a legacy pipeline to dbt/SQLMesh with a reconciliation harness.
Months 6-12
Take ownership of orchestration, CI/CD, lineage, and a semantic layer - become the platform owner, not the pipeline writer. That's the role at the top of the range.
Gear for this job
As an Amazon Associate, PayCrunch earns from qualifying purchases. Links to books and tools are for the job on this page; we only recommend what we’d use in the work.
Same live O’Reilly Jan 2024 already on data-analyst / data-engineer / business-intelligence-analyst / data-architect. This page’s first play is Ship pipelines faster with AI-generated dbt models and tests and Month 1 is Turn on dbt Copilot. Not official dbt Labs cert and not CompTIA Data+.
Next steps for an ETL Developer
Some links below are affiliate or partner links. PayCrunch may earn a commission if you enroll or subscribe through them, at no extra cost to you. Wage figures on this page still come from the Bureau of Labor Statistics, not from these programs.
ETL Developer work is specific enough that a stamped 'check out these courses' block would be noise. BLS files this work as Database Administrators (SOC 15-1242). O*NET Job Zone 4 is typical: a bachelor's degree, so the honest next credential is a professional certificate or bachelor's-level coursework — not a random catalog dump.
The occupation's listed knowledge areas include Telecommunications and Engineering and Technology; the links search those subjects, not a generic 'career courses' list.
ETL Developers in this dataset list AJAX among the tools in use, so a program that names that stack is a better fit than a survey course.
Coursera search for telecommunications — a professional certificate or bachelor's-level coursework that lines up with computing, not a generic professional-development aisle.
FlexJobs screens remote, hybrid, freelance, and flexible listings so you are not wading through unverified ads. This is a job-board search for ETL Developer work, not a claim that they list a counted SOC 15-1242 inventory.
Write an ETL Developer resume, or one aimed at Software Developers, instead of a blank template. Resume Now is a resume builder; we are not claiming a counted template set for this SOC.
An ETL Developer resume that names the actual tasks on this page, or the step-up title Software Developers, beats a blank template when you apply.
What ETL Developers earn by state
These are the Bureau of Labor Statistics’ own figures for Database Administrators, state by state — not a cost-of-living adjustment applied to the national number. Only states employing at least 500 people in the occupation are shown, because a state median drawn from a handful of workers is noise rather than a signal.
Utah
$135,750
highest of them · +30% vs the national median
Kentucky
$87,390
lowest of the 33 states and D.C. that qualify · -16% vs the national median
The same job pays $48,360 more a year at the median in Utah than in Kentucky — 55% higher. That gap is what the Bureau measured, before any question of what it costs to live in either place. The top-of-range figure quoted at the head of this page, $178,580, is a different statistic in a different place: it is the 90th-percentile wage in New York. The state that pays the typical worker most and the state where the best-paid go highest are not always the same one.
Source: U.S. Bureau of Labor Statistics, Occupational Employment and Wage Statistics, May 2025, SOC 15-1242. 33 states and D.C. clear the 500-employee reporting floor for this occupation; those below it are left out rather than shown with a wide error band.
Free data. Use any of it.
PayCrunch publishes verified, BLS-sourced salary + AI-playbook data on 1,000+ professions — free, no signup.
Honestly, the narrow version of the job - hand-writing routine transforms from a ticket - is the most exposed work in data, and AI does it well. But 'move data correctly and reliably' is not the same as 'write SQL.' Someone must design the models, guarantee the numbers are right, own cost and governance, and answer when a pipeline breaks. Developers who climb into platform, quality, and architecture roles are more valuable because of AI; those who only write transforms are the ones at risk.
Is it safe to paste our data or SQL into ChatGPT?
Not real data, credentials, or connection strings. Use synthetic column names and sample rows, or an enterprise AI instance your company has approved. The logic and schema shape are usually enough for the AI to help - you don't need to expose actual customer records to get a good model or test suite.
AI writes the SQL now - is it even worth deepening my skills?
More than ever, but in a different direction. The scarce skill is no longer typing joins; it's knowing when the AI's join is wrong, designing a model that won't fan out, catching a silent type-cast bug, and architecting for cost and reliability. That judgment is what you should deepen - AI handles the syntax, you own correctness.
How does AI actually increase an ETL developer's pay?
By freeing you to move up the value chain. When AI writes the boilerplate, you spend that time on data quality, cost optimization, migrations, and platform ownership - the work that's visibly worth more. A measurable warehouse-bill cut or a led migration is exactly the evidence that justifies a jump toward the $179k top of the range.
Which AI skill should I build first?
AI-assisted dbt (or your warehouse's SQL) with rigorous test generation. It's the fastest productivity win and it forces the right habit: never ship a machine-written transform without tests that prove the numbers. That discipline is what distinguishes an engineer the business trusts from one it double-checks.
Methodology & sources
Salary (median, 10th, top of the range) — U.S. Bureau of Labor Statistics, OEWS.
By state — the Bureau of Labor Statistics’ own state medians, limited to states employing at least 500 people in the occupation. No cost-of-living arithmetic is applied to a wage anywhere on this page.
The plays — PayCrunch's own step-by-step guidance using publicly available AI tools. Tool names/URLs are real and current as of August 2026; prompts written to work as-is. Verify any professional output before relying on it.