Why prompt engineers stall without a second qualification
$234,980estimated top of the range · middle $120,000 / yr
AI is creating this demand
Prompt Engineers in the United States earn a median of $120,000 a year. Pay starts near $72,000. The top of the range is estimated at $234,980. The Bureau of Labor Statistics does not publish a separate wage series for this exact title, so this figure is derived from the closest occupation it does track and is labelled an estimate.
Source: PayCrunch estimate. Last checked 9 September 2026.
Entry level
$72,000
Top-end estimate
$234,980
Education
Bachelor's degree in CS or Linguistics
Wages — PayCrunch estimate. The Bureau of Labor Statistics does not publish a separate wage series for Prompt Engineer; figures are derived from the closest occupation it does track and are labelled as estimates. AI-impact rating is PayCrunch's editorial assessment. Updated September 2026.
🆕 New & Trending AI Tools for Prompt EngineerReviewed September 2026
We track new AI-tool launches every week and refresh this list — here’s what’s gaining traction for Prompt Engineer work right now.
Claude CodeNEWFree / usage-based
Terminal coding agent that reads your repo, runs tests, and ships multi-file changes.
How a Prompt Engineer uses it: describe a feature and let it implement and test it across the codebase
OpenAI CodexNEWIncl. w/ ChatGPT plans
Agent that runs longer, deterministic multi-step coding jobs on its own.
How a Prompt Engineer uses it: delegate a well-defined build or migration and review the finished result
WindsurfNEWFree / $15 mo
Agentic IDE that keeps context across a whole project.
How a Prompt Engineer uses it: make large, coordinated changes without losing track of the codebase
AWS KiroNEWPreview / see site
Spec-driven coding agent that turns written specs into working code.
How a Prompt Engineer uses it: write the spec first and let it build to that spec
NotebookLMNEWFree / $7.99 mo
Google tool that answers questions grounded only in the documents you give it — with citations.
How a Prompt Engineer uses it: load your own manuals, policies, or PDFs and ask questions that stay accurate to the source
CursorFree / $20 mo
AI-native code editor that edits across an entire project.
How a Prompt Engineer uses it: describe a change in plain English and let it rewrite and refactor whole files
GitHub Copilot (Agent Mode)$10–19 mo
AI pair-programmer built into VS Code and GitHub that now completes multi-step tasks.
How a Prompt Engineer uses it: hand off a task and have it plan, edit multiple files, and open a pull request
ChatGPTFree / $20 mo
The most-used AI assistant — writing, analysis, research, and images from a plain-language chat.
How a Prompt Engineer uses it: draft emails and documents, summarize long files, and get instant answers to on-the-job questions
ClaudeFree / $20 mo
AI assistant known for careful writing, long-document analysis, and coding.
How a Prompt Engineer uses it: analyze big reports or spreadsheets and turn messy notes into clean, finished writing
The wording the product actually ships
A prompt engineer writes and tests prompts for a product team. The prompt is the instruction a model sees when a feature runs: the wording that tells it what job it is doing, what it should refuse, and what shape the answer should take. You do not own the whole product. You own the language that makes this part of the product useful, and you own the habit of checking that language against cases the team cares about.
The day sits next to engineers and a product manager. Someone has a feature that needs language. You draft it. You run it against examples. You watch where the answer wanders, gets vague, or misses the task. You change the wording and you run it again. Then you write down what you changed, because a prompt that only lives in a chat window will be lost by Friday. The craft is repetitive on purpose. A clever sentence that fails the second example is not done.
Writing, then testing, then rewriting
Writing starts with the task in the user's words, not in the model's. What should a person be able to do after this feature answers? What must the answer include? What should it leave out? You turn those decisions into a prompt the model can follow, and into a note the rest of the team can read. Short prompts and long ones both fail when the task is muddy. Your first move is often to sharpen the task with the product manager before you polish a sentence.
Testing is the half of the job outsiders skip. You keep a set of examples: ordinary requests, odd requests, and the cases where a bad answer would embarrass the product. You run the prompt. You compare what came back with what the team agreed to want. You look for answers that are confident and wrong, answers that ignore a constraint, and answers that drift into a tone the brand will not accept. Then you revise. A prompt engineer who only demos the happy path will ship a feature that breaks the first week a real customer touches it.
Rewriting is where judgment shows. You change one thing at a time when you can, so you know what helped. You save the version that failed and the version that held. You talk with the engineer about what belongs in the prompt and what belongs in code around it. Some failures are wording. Some are the information the model was given. Some are a product decision nobody has made. You are useful when you can tell those apart and say so without drama. The library of prompts becomes a product asset: named, versioned, and tied to the examples that justify them.
How people get good at the craft
There is no licence for this title. People arrive from writing, from support, from linguistics, from engineering, from research, and from product roles where they were already the person who edited the model's instructions. A degree can help you get the conversation. A folder of work helps you get the job. The folder should show prompts you wrote, the examples you tested, and what you changed after a failure. Strip anything confidential. If the only artifact is a screenshot of a fun chat, you have a hobby. If the artifact shows a task, a failure, and a revision, you have the beginning of the craft.
Practice on a product-shaped problem, not on a party trick. Pick a task a real team might ship: summarizing a support thread into a reply a human will edit, turning a messy note into a structured record, drafting a first answer that a specialist will review. Write the prompt. Build a handful of examples, including ones designed to tempt a sloppy answer. Revise until you can explain why the wording is the way it is. Then ask an engineer or a writer to break it. The breaks are the education.
Keep a short log beside the library. For each prompt, note the task, the date of the last revision, and the example that most recently failed. The log is how a teammate can change your wording without guessing. It is also how you remember, months later, why a sentence looks odd. Teams that skip the log rediscover the same failure in front of a customer. Teams that keep it can review a change in an afternoon. Treat the log as part of the prompt, not as homework you will do if the week is quiet.
You will also sit with the people who worry about harm. A product team needs language that stays inside what the feature is allowed to do. Your role there is to make the prompt match decisions the company has already made, and to surface cases where the model slips past those decisions. You bring the examples. You do not freelance a policy in the wording. When a case is unclear, you stop and ask the person who owns the decision. That pause is part of the craft, and hiring teams can hear it when you tell the story of a revision you refused to ship alone.
Learn enough about how the product is built to be a good partner. You do not need to train a model from scratch to be hired. You do need to know where a prompt sits in the feature, how an example is stored, and how a change gets reviewed before it reaches customers. Read the team's existing prompts before you propose a new voice. Consistency across a product matters more than a single brilliant paragraph. The people who grow in this work are the ones other teams ask for when a feature sounds clever and behaves badly.
What a hiring loop is listening for
A loop usually includes the product manager, an engineer, and sometimes a designer or a researcher. They want to hear you revise. Bring a prompt you are willing to watch fail in the room. Walk through an example that broke it and the wording you tried next. Explain a constraint you put in on purpose. Explain one you removed because it made ordinary answers worse. That story is more persuasive than a claim that you are good with language.
Expect a live problem. They may describe a feature and ask you to draft a prompt and a few examples you would test. Think aloud. Separate the product decision from the sentence. Ask who the user is and what a bad answer costs. Hiring teams trust a person who names a missing fact. They worry about a person who performs certainty and then cannot say how they would notice a failure after launch. Ask them how prompts are reviewed, who can ship a change, and whether you would sit with one squad or float across many. The answer tells you whether the title is a real seat.
If you are early, say so, and show the practice folder. A support agent who has been repairing bad model replies for a year has a kind of experience a classroom cannot fake. A writer who has never looked at a failure case will struggle. Be honest about which one you are. Then show the revision habit. That habit is the hire.
Beyond the first library
The early job is one feature, one library, one squad. You learn the product's voice and the engineer's constraints. The next step is often a broader surface: several features that must sound like one product, or a platform other teams use when they add a model to their own work. Senior people spend more time on the examples and the review path, and less time on a single sentence. They teach other writers and engineers how to test a change. They decide when a prompt should be retired because the product decision changed.
Some people move toward research, evaluation design, or a product manager seat, because they have spent a year watching what language can and cannot fix. Some stay and become the person who knows every prompt the company ships. Both are legitimate. What ages badly is a private style with no examples behind it. Keep the library legible. When you look at the next role, ask how failure is found after launch, and whether you will be in that conversation. A prompt engineer who is excluded from the failures is being asked to decorate. A prompt engineer who is handed the failures is being asked to practice the job.
Estimated pay for prompt work
Read these as estimates
PayCrunch built these estimates for the title itself. The Bureau of Labor Statistics does not publish a separate wage series for this exact title. The dollars describe someone who writes and tests prompts for a product team. They do not assign a figure to a state.
The entry estimate is $72,000. The median estimate is $120,000. The top estimate is $234,980. From entry to the median is $48,000. From the median to the top is $114,980. A first seat, or a move in from an adjacent writing or support role, can sit near $72,000. The median of $120,000 describes a prompt engineer who already owns a library, tests it against real cases, and ships revisions with a squad. The top of $234,980 belongs with range: several products, a review path other people follow, and a record of failures caught before customers did. The $114,980 above the median is a long step. Treat it as the shape of the estimate, not as next year's raise.
Say the word estimate when you repeat $72,000, $120,000, or $234,980. A recruiter who drops one of those numbers into a city name is adding a claim the estimate does not make. Use $48,000 as the distance from entry to the middle. Use $114,980 as the distance from the middle to the top. Those two distances do different work in a negotiation, and they should not be blended into a single hopeful leap.
Three figures and a conversation
Before you answer an offer, write the base beside $72,000, $120,000, and $234,980. If the base is near $72,000, you are looking at the entry estimate. The median is $48,000 higher. If you already have a library and a testing habit a team has used, name that gap and tie it to the feature you would own. If you are new, a year near $72,000 can be the right trade for a squad that will let you see failures and revise in public. Decide which story is true before you perform the other one.
If the offer is near $120,000, it matches the median estimate. The remaining distance is $114,980, up to $234,980. Ask whether the role is one feature or a practice other teams will copy. Ask who reviews a prompt before it ships, and whether you are expected to build that review. Scope of that kind is what the top of the estimate is pointing at. A single feature with a careful product manager belongs nearer the median. Put bonus and equity in separate sentences so they do not blur the base.
Inside a job, bring the same three markers to a review, along with the prompts you shipped and the failures you caught. The estimates are PayCrunch figures because the Bureau of Labor Statistics does not publish a separate wage series for this exact title. Your evidence is concrete: wording a product team could maintain, examples that predicted real mistakes, and revisions that made the feature safer to ship. That record is how you talk about $72,000, $120,000, and $234,980 without turning the conversation into a mood.
The top of Prompt Engineer pay — and how to get there with AI
$234,980top-end estimate for Prompt Engineer
PayCrunch estimate - derived from the closest occupation BLS tracks (Software Developers, 15-1252). This figure is PayCrunch’s estimate, not a Bureau of Labor Statistics published wage for this exact title.
And the role it leads to — Computer Hardware Engineers — reaches $281,210 in California.
$72,000entry$120,000middle$234,980top end
Prompt engineers near the top of the range are rarely the ones with the neatest wording; they hold a second examined qualification in measurement, security review or system cost, so their claim that a build behaves as specified can be checked.
Wording alone is cheap to copy and impossible to defend at review time. The part of this job that keeps its value looks like engineering: storing, retrieving and manipulating the data needed to analyse what a system can and cannot do, preparing reports on project specifications and status, and weighing reporting formats, running costs and security needs before a configuration is chosen. Assistants such as Claude or ChatGPT write the candidate wording faster than any person can. Nobody has automated the qualification that says your evaluation method holds up, and that is the piece employers buy.
Your playbook, by where you are now
Just startingMake your work measurable before you make it clever
Keep every version of every prompt, its inputs and its outputs in Airtable so a change can be traced to a result.
Write a fixed set of hard cases before you start tuning, and score against that set rather than against your own impression.
Record how long each build takes and how much compute it consumed on Amazon Elastic Compute Cloud EC2, because cost is part of the specification.
Write a short status report for each project in the format the people funding it actually read.
Take a graded course in experimental design or applied statistics and sit the examination, not just the lectures.
What proves it: A scored evaluation set with a written method somebody else could rerun.
Realistic span: the first eighteen months
A few years inAdd the qualification the hiring bar assumes
Choose one adjacent field, security review, cloud architecture or data engineering, and finish its certification inside a year.
Apply it immediately: run a written security and data-handling review of one system before it reaches users.
Move the stored evaluation data into Amazon DynamoDB or Amazon Redshift so results survive the person who produced them.
Build the pipeline in Cursor or with GitHub Copilot, then read every generated line before it becomes part of a shipped configuration.
Monitor a live system against its written specification and publish the drift monthly instead of when somebody complains.
What proves it: A certificate in the adjacent field plus one review document that changed a launch decision.
Realistic span: years two through five
ExperiencedSet the standard others are measured against
Write the organisation's evaluation standard and make it the gate a build has to clear before release.
Supervise the programmers and technicians doing this work, assigning cases by risk rather than by who is free.
Train users on each modified system yourself, since adoption is where most of these projects quietly fail.
Evaluate reporting formats, running costs and security requirements before any vendor commitment, and put the recommendation in writing.
Move toward hardware and systems engineering scope, where California employers price this depth highest.
What proves it: An evaluation standard adopted beyond your own team, with your name on its revision history.
Realistic span: year five onward
The next 90 days
In the next ninety days, build one honest evaluation set. Take the system you touch most, collect forty real cases including the ones that embarrass it, write down what a correct output looks like for each, and score your current configuration against them. Store the cases, the outputs and the scores in Airtable so the record outlives the sprint. Then enrol in a qualification that teaches you why that scoring is sound, statistics, security review or cloud architecture, and finish it. Wording is a skill anybody can claim; a scored method plus an examined credential is a claim somebody can verify.
Wage figures: PayCrunch estimate. The playbook is PayCrunch editorial guidance, not a guarantee of pay or placement.
Every figure is the national median from the U.S. Bureau of Labor Statistics (OEWS) shown on that role’s own page.
Never used AI before? Start here (2 minutes).
Open an eval tool before you open a prompt editor. The habit that separates top prompt engineers from the rest is measuring, not guessing. Sign up for Braintrust or install open-source Promptfoo, and build a 20-example test set for the task you're working on. Now every prompt change is scored against real cases instead of a single lucky demo.
For the craft itself, read the official prompting guides from Anthropic and OpenAI, and the DAIR.AI Prompt Engineering Guide — all free. Practice in the Anthropic Console Workbench and the OpenAI Playground where you can compare models side by side. AI writes the draft prompt; your eval set tells you the truth.
The one rule, forever: Treat every prompt as an attack surface. Never paste secrets, API keys, or real user PII into prompts or logs; assume untrusted input can carry prompt-injection and test for it; and never ship an LLM feature to production on vibes — a change that isn't measured against an eval set can silently regress. In regulated domains, keep a human in the loop and log every model decision.
The plays — exact steps, exact prompts
Do these in order. Each one is copy-paste ready. You do not need to know anything about AI going in.
1
Build an eval suite so you optimize with data, not vibes
Why this pays: The single skill that moves a prompt engineer from $120,000 to $185k is proving reliability. Teams pay for someone who can raise a feature's accuracy from 82% to 96% and show the number — evals are how you demonstrate that value.
BraintrustPromptfooLangSmith
1
Curate a golden dataset of 30-100 real input/output pairs for your task in Braintrust (or a YAML config in Promptfoo). Include the hard edge cases and past failures — that's where regressions hide.
2
Have an LLM help you design graders (LLM-as-judge plus deterministic checks) for the dimensions that matter.
Copy-paste this prompt
I'm building an eval for an LLM feature that [extracts structured invoice fields from raw email text]. Propose a scoring rubric with 4-6 criteria (e.g. field accuracy, format validity, hallucination rate, refusal-when-uncertain). For each criterion, tell me whether it should be a deterministic check or an LLM-as-judge grader, and write the judge prompt for each LLM-graded one. Make the judge prompts strict and specific.
Use LLM-as-judge for fuzzy quality and deterministic asserts for anything checkable (JSON valid, field matches). Spot-check the judge against your own labels — a bad grader gives confident wrong scores.
3
Run every prompt or model change through the suite and gate deploys on the score. Track results over time so you can prove the trend line, not just today's demo.
What you'll haveA reliability number you can move and defend — the artifact that gets a prompt engineer trusted with production systems and paid like an engineer.
2
Version, observe, and A/B test prompts in production
Why this pays: Prompts that live in code comments don't scale. The engineer who versions prompts, watches production traffic, and ships measured improvements owns the feature — and owning a revenue feature is how you reach the top band.
PromptLayerLangfuseHelicone
1
Move prompts out of hard-coded strings into PromptLayer or Langfuse so each has a version, a changelog, and metrics — and you can roll back a bad prompt without a code deploy.
2
Instrument production with Helicone or Langfuse to capture latency, cost per request, and token usage per prompt version, then mine the traces for failures.
Copy-paste this prompt
Here are 15 production traces where users thumbs-downed the response [paste sanitized traces]. Cluster them by root cause (retrieval miss, ambiguous instruction, format failure, hallucination, tone). For each cluster, tell me the single prompt or pipeline change most likely to fix it, and how I'd measure whether it worked.
Strip all PII from traces before pasting into any external tool. Fixes are hypotheses — confirm each against the eval suite before shipping.
3
Run an A/B or shadow test between prompt versions on live traffic and promote the winner on evidence, not intuition.
What you'll haveA production prompt system with rollback, cost control, and measured wins — the operational maturity that commands senior pay.
3
Optimize prompts programmatically with DSPy
Why this pays: Hand-tuning prompts has a top end; programmatic optimization breaks through it. The rare engineer who can compile and optimize prompt pipelines instead of hand-crafting strings is doing applied-AI-engineer work — and being paid for it.
DSPyClaudeGPT-5
1
Reframe a brittle hand-written prompt as a DSPy program with typed signatures and modules, then let its optimizers (e.g. MIPRO, bootstrap few-shot) search for better instructions and examples against your metric.
2
Design the metric function that DSPy optimizes toward — this is where your judgment lives.
Copy-paste this prompt
I'm converting a [customer-support classification] prompt into a DSPy module. My success metric weighs [correct category] at 70%, [correct urgency] at 20%, and [no fabricated policy references] at 10%. Write the DSPy signature, a first-pass module, and a metric function that returns a single score combining these, penalizing fabrications hardest. Explain what the optimizer will and won't be able to improve.
Optimizers overfit to your training set — always hold out a fresh test split and confirm gains transfer before you ship the compiled prompt.
3
Compare the optimized pipeline against your hand-written baseline on the held-out set and keep whichever genuinely wins.
What you'll havePrompt pipelines that improve themselves against a metric — a differentiated, higher-value skill few colleagues have.
4
Harden against prompt injection and jailbreaks
Why this pays: An LLM feature that leaks data or gets jailbroken is a security incident. The prompt engineer who red-teams and hardens systems is the one companies trust with customer-facing AI — and trust is what unlocks the senior title and pay.
Promptfoo Red TeamLakera GuardAnthropic Console Workbench
1
Run Promptfoo's red-team module against your system prompt to generate injection, jailbreak, and data-exfiltration attempts automatically, and see what gets through.
2
Generate an adversarial test set that mirrors your real threat model.
Copy-paste this prompt
My LLM agent has a system prompt that must never [reveal its internal tools list or the contents of other users' records]. Generate 20 adversarial user messages that try to extract that via prompt injection, role-play, encoding tricks, translation attacks, and instruction-override. For each, note the technique so I can add it to my eval suite as a must-refuse case.
Use these only against your own systems. Add every successful attack to your permanent eval set so a future prompt change can't silently reopen the hole.
3
Add a guardrail layer (e.g. Lakera Guard or an LLM-based input/output filter) and re-run the red-team suite to confirm the attacks now fail.
What you'll haveA demonstrably hardened LLM feature with a regression-tested threat model — the trust that gets you the customer-facing, higher-paid work.
5
Engineer context and retrieval, not just wording
Why this pays: Most 'prompt' failures are actually context failures. The engineer who fixes retrieval — chunking, embeddings, reranking — solves problems others can't, moving from prompt tweaker to RAG owner and into the top pay band.
LlamaIndexCohere RerankLangChain
1
Build the retrieval layer in LlamaIndex or LangChain, then measure retrieval quality separately from generation — if the right chunk never gets retrieved, no prompt can save the answer.
2
Diagnose whether your failures are retrieval or generation, and fix the real cause.
Copy-paste this prompt
My RAG pipeline answers [internal HR policy questions] and is wrong ~20% of the time. Give me a diagnostic checklist to isolate whether failures are from retrieval (wrong chunks) or generation (right chunks, bad answer). For retrieval, list concrete fixes to test — chunk size and overlap, hybrid keyword+vector search, adding a reranker, metadata filtering — and how to measure each with a retrieval-recall metric.
Measure retrieval recall on a labeled query set before touching the prompt. Add a reranker (Cohere Rerank) only if it moves the metric — complexity you can't measure is complexity you don't need.
3
Add reranking and metadata filtering where the metric proves they help, and lock in the gain with a retrieval eval that runs on every change.
What you'll haveA retrieval system tuned by measurement — solving the failures that pure prompting can't, and owning the whole RAG stack.
6
Build agents and tool-use that actually finish the task
Why this pays: Single-shot prompting is commoditizing; reliable multi-step agents are not. The engineer who can make an agent complete a real workflow end to end owns the hardest, most valuable AI work — the frontier of the pay band.
LangGraphOpenAI Agents SDKAnthropic tool use
1
Model the workflow as an explicit graph in LangGraph (or the OpenAI Agents SDK) with defined tools, state, and stopping conditions — not one mega-prompt hoping the model figures it out.
2
Design tight, unambiguous tool definitions and the decision logic for when to call each.
Copy-paste this prompt
I'm building an agent that [triages inbound support tickets: looks up the customer, checks order status, drafts a reply, and escalates if refund > $200]. Write clear tool schemas (name, description, JSON parameters) for each tool, the system prompt that governs when to call which, and the guardrail that forces escalation on the refund threshold. Then list the 8 failure modes I should build eval cases for.
Constrain tools narrowly and make the model's authority explicit — an agent that can act needs hard limits. Eval multi-step trajectories, not just final answers, so you catch mid-run errors.
3
Trace and eval full agent trajectories (LangSmith or Langfuse) so you catch the step where a multi-step run goes wrong, and gate releases on task-completion rate.
What you'll haveAgents that reliably complete real workflows, measured end to end — the highest-value LLM engineering work there is.
Your 12-month sequence to the top of the range
How the plays above stack into a path from median pay toward the $185,000 tier.
Month 1
Build your first golden eval set in Braintrust or Promptfoo and make 'no prompt change ships without a score' your default. Learn the Anthropic and OpenAI prompting guides cold.
Months 2-3
Move prompts into a versioned store (PromptLayer/Langfuse), instrument production observability, and start mining traces for failure clusters.
Months 3-6
Own a real RAG feature: measure retrieval separately, add reranking where it helps, and drive an accuracy number up with evidence.
Months 6-9
Red-team your systems with Promptfoo, add guardrails, and learn DSPy to optimize a pipeline programmatically instead of by hand.
Months 9-12
Ship a measured multi-step agent with trajectory evals and task-completion gating — the portfolio piece that justifies an applied-AI-engineer title.
Year 2
Position as the eval-and-reliability owner for LLM features across teams — the scarce role that pays at the top of the band.
Next steps for a Prompt Engineer
Some links below are affiliate or partner links. PayCrunch may earn a commission if you enroll or subscribe through them, at no extra cost to you. Wage figures on this page still come from the Bureau of Labor Statistics, not from these programs.
Prompt Engineer work is specific enough that a stamped 'check out these courses' block would be noise. BLS files this work as Software Developers (SOC 15-1252). O*NET Job Zone 4 is typical: a bachelor's degree, so the honest next credential is a professional certificate or bachelor's-level coursework — not a random catalog dump.
Prompt Engineers in this dataset list AJAX among the tools in use, so a program that names that stack is a better fit than a survey course.
The next title this dataset points at is Computer Hardware Engineers; a credential aimed that way is a clearer step than another year in the same seat.
Coursera search for computer science — a professional certificate or bachelor's-level coursework that lines up with computing, not a generic professional-development aisle.
FlexJobs screens remote, hybrid, freelance, and flexible listings so you are not wading through unverified ads. This is a job-board search for Prompt Engineer work, not a claim that they list a counted SOC 15-1252 inventory.
Write a Prompt Engineer resume, or one aimed at Computer Hardware Engineers, instead of a blank template. Resume Now is a resume builder; we are not claiming a counted template set for this SOC.
A Prompt Engineer resume that names the actual tasks on this page, or the step-up title Computer Hardware Engineers, beats a blank template when you apply.
What Prompt Engineers earn by state
This page does not show a state table, and the reason is worth stating: the Bureau of Labor Statistics does not publish a separate wage series for this job title, so there are no official state figures to show. Scaling the national median by a cost-of-living index would produce a number for every state, but it would be an estimate of living costs wearing a wage’s clothes, and PayCrunch would rather show you nothing than that.
What the national figures say: pay starts near $72,000, the median is $120,000, and the top of the range is $234,980. Those national figures are a PayCrunch estimate, not a Bureau of Labor Statistics published wage for this exact title.
The naive version of the job — hand-crafting clever wording — is already commoditizing, because models get better at understanding plain instructions every release. But someone still has to define what 'correct' means, build the eval sets, engineer retrieval and agents, and harden against attacks. That work is growing, not shrinking. The title may fold into 'AI engineer,' but the skills are exactly what stays valuable.
Is prompt engineering a real, durable career?
The durable version is. If your whole value is a folder of prompt templates, that's fragile. If your value is eval-driven reliability, RAG, agents, and security — the systems that make LLMs work in production — that's applied AI engineering, and demand for it is strong. Aim your growth at the systems, not the wording.
How does building evals actually raise my pay?
Because it converts 'the AI feels better' into 'accuracy went from 82% to 96%, here's the graph.' Teams pay senior rates for someone who can reliably improve a production metric and prevent regressions. Evals are the proof-of-value that turns a $120,000 prompt writer into a $185k engineer who owns a feature.
Which model should I build on — GPT, Claude, or Gemini?
All of them, and stay model-agnostic. Your eval suite lets you swap models and measure the tradeoff in accuracy, latency, and cost objectively. The engineers who get stuck are the ones whose prompts are hand-fit to one model; the ones who thrive treat the model as a swappable component behind a tested interface.
Is it safe to paste production data into prompt tools?
Not raw. Strip PII, secrets, and API keys before any input touches an external tool or a log, and assume untrusted user input may carry prompt injection. Use enterprise tiers with no-training terms for anything sensitive, add guardrails on customer-facing systems, and keep a human in the loop for regulated decisions.
Methodology & sources
Salary (median, 10th, top of the range) — U.S. Bureau of Labor Statistics, OEWS.
By state — the Bureau of Labor Statistics’ own state medians, limited to states employing at least 500 people in the occupation. No cost-of-living arithmetic is applied to a wage anywhere on this page.
The plays — PayCrunch's own step-by-step guidance using publicly available AI tools. Tool names/URLs are real and current as of August 2026; prompts written to work as-is. Verify any professional output before relying on it.