Job family
Assessing quality engineering and test
Roles paid to find the defect before the customer does, and to build the harness that keeps finding it.
Quality engineering has quietly become the most consequential technical function that nobody screens properly. The BLS counts 187,600 US quality assurance analysts and testers inside a 1.9 million-strong occupation group, paid a median of $104,300 — roughly $32,000 below the developers whose work they gate. The pay gap tells you where the hiring attention goes. It does not tell you where the risk is: an organisation shipping code faster than it can reason about it needs the test function to be sharper than the development function, not cheaper.
The distinguishing skill in this family is subtractive, and that is what makes it hard to see in an interview. Anyone can list test cases. A strong quality engineer decides what not to test, argues for the three scenarios that carry almost all the risk, and writes a bug report that a developer can reproduce on the first attempt without a follow-up conversation. The failure modes are specific and expensive: a suite so flaky that the team learns to ignore red builds, coverage metrics that climb while escaped defects climb with them, automation written against implementation details so that every refactor breaks a hundred tests. None of these are visible from a candidate who can write a passing Selenium script.
The pressure on screening here is sharper than in development. The 2025 DORA research found 90 percent of technology professionals using AI at work and more than 80 percent reporting a productivity gain, while 30 percent said they had little or no trust in AI-generated code; Stack Overflow's 2025 survey found 45 percent of respondents saying that debugging AI-generated code takes more time, not less. Both statements describe a world where the volume of plausible-looking code needing verification is rising and the reliability of that code is not. Meanwhile the standard QA screen — write a test for this function, describe the difference between smoke and regression testing — is exactly the kind of self-contained prompt that a model answers perfectly.
The design that survives is a monitored sandbox containing a small application with a real, non-obvious defect, plus a follow-up interview about the candidate's own submission. What did you decide not to cover, and why? Which of your assertions would still fail if the bug were fixed a different way? A candidate who generated their suite cannot answer the second question; a candidate who reasoned about risk answers it immediately. That is the entire difference, and it is invisible on a pass/fail test run.
Why this work can be assessed
Testing is a written artefact discipline — a test plan, a bug report, an automated suite — all of which can be produced in a monitored sandbox and then defended in a spoken interview about the candidate's own choices of what not to test.
Roles in this family
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Software Developers, Quality Assurance Analysts, and Testers, 2025, https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm
- Stack Overflow, 2025 Developer Survey, AI section, 2025, https://survey.stackoverflow.co/2025/ai
- Google Cloud, Announcing the 2025 DORA Report: State of AI-assisted Software Development, 24 September 2025, https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
Hiring for one of these? We build the assessment for the specific role, run it under your brand, and return a ranked list with the evidence behind every score.
Book a walkthrough