Get started

Job family

Assessing machine learning and ai engineering

Roles that put a model into production and stay responsible for what it does there.

This is the hardest family in technical hiring to define, and the labour statistics show why: the closest occupational code, computer and information research scientists, counts 38,600 US jobs, which is obviously far fewer people than are employed doing machine learning work. The demand shows up first in postings. Lightcast's data for the Stanford AI Index recorded US postings mentioning generative AI as a skill rising roughly fourfold in a single year, to more than 66,000 in 2024, with large language modelling quadrupling to 20,000. The consequence for a hiring manager is that the title tells you almost nothing. "AI Engineer" in one company means fine-tuning and evaluation; in another it means calling a hosted API behind a feature flag.

The candidates are correspondingly hard to read. Model architecture knowledge is now widely available and shallowly held, and a large share of applicants can discuss transformers fluently without ever having shipped something that degraded in production. The separation between a strong and a weak hire in this family is almost entirely about evaluation and failure: whether they build an honest eval set before they build the model, whether they can tell you what their offline metric fails to capture, whether they notice label leakage, whether they can say out loud that the model is not the right tool for this problem. The expensive failure is not a bad model — it is a plausible model deployed with no way to detect that it has stopped working.

Screening in this family was already weak and is now largely broken. The standard artefacts are a Kaggle-style notebook take-home and a set of textbook questions on bias and variance, and both are exactly the shape of task that current models complete to a high standard unsupervised. Worse, this is the one family where candidates can reasonably argue that using AI is the job, which makes a naive ban both unenforceable and slightly absurd.

The productive answer is to stop testing whether they can produce a model and start testing whether they can break one. A monitored sandbox with a pre-existing model and a dataset containing a specific pathology — a leaked feature, a distribution shift between train and test splits, a metric chosen to flatter — gives the candidate something to diagnose rather than generate, and the follow-up interview about their own diagnosis is where the reasoning either holds up or does not. Buyers should hold this family's fit at medium rather than high, and be told plainly that a timed exercise reads judgment about a known problem well and predicts nothing about whether someone can carry an eighteen- month research programme.

Why this work can be assessed

Much of the job is code and written reasoning about evaluation, both producible in a sandbox — though the longest and most valuable feedback loops in this family run for months and no timed exercise reproduces them.

Roles in this family

AI EngineerThis is the one role where a candidate can correctly argue that using AI is the job, which makes the standard integrity posture in…Machine Learning EngineerThe notebook take-home supplies a target variable, a metric and a split, which are the three decisions this role is paid to make, …MLOps EngineerThe whiteboard architecture question is answered from a reference diagram, and reference diagrams are exactly what a model produce…

Sources

Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.

  1. US Bureau of Labor Statistics, Occupational Outlook Handbook, Computer and Information Research Scientists, 2025, https://www.bls.gov/ooh/computer-and-information-technology/computer-and-information-research-scientists.htm
  2. Stanford HAI, AI Index Report 2025, labour market data by Lightcast, 2025, https://lightcast.io/resources/research/stanford-ai-index-2025
  3. US Bureau of Labor Statistics, Occupational Outlook Handbook, Data Scientists, 2025, https://www.bls.gov/ooh/math/data-scientists.htm

Hiring for one of these? We build the assessment for the specific role, run it under your brand, and return a ranked list with the evidence behind every score.

Book a walkthrough