Job family
Assessing platform, devops and site reliability
Roles that own the systems everyone else deploys onto, and the pager that rings when those systems fail.
Read the BLS projection for systems administrators literally and you would conclude this family is shrinking: 323,600 jobs in 2025, down 4 percent by 2035, with every one of its roughly 13,400 annual openings attributed to replacement rather than growth. Read it against the DORA finding that 90 percent of surveyed organisations now run at least one internal platform, and a different picture emerges. The headcount is not disappearing; it is being re-titled. The person who used to patch servers is now expected to write Terraform, own an SLO, run a blameless post-incident review and treat other engineers as customers of a product they maintain.
That re-titling is precisely why hiring in this family goes wrong. The job description gets written as a tool list — Kubernetes, Terraform, one of the three big clouds, a CI system — and the interview then tests recall of those tools. Tool recall is the least durable thing about the role. What actually separates a strong platform or reliability hire is behaviour under partial information: forming a hypothesis from a dashboard that is itself possibly lying, choosing to mitigate before understanding root cause, knowing when to stop debugging and roll back, and reasoning about blast radius before touching production. The costly failures are judgment failures — a "quick fix" applied at 2am that takes out a second region, a runbook nobody could follow, an alert threshold set so tight the team stopped reading the channel.
This family is less exposed to AI-assisted cheating than pure development work, for a structural reason worth stating: an assistant can write the Terraform, but it cannot tell the candidate what is actually broken in a system it cannot see. That makes an incident simulation unusually robust. It also means the traditional screen fails for a different reason than elsewhere — not because candidates can fake it, but because a whiteboard question about container networking never observed anyone making a decision under pressure in the first place.
The design follows from that. Put the candidate into a monitored sandbox with a degraded service and partial, slightly misleading telemetry; watch what they check first and how they narrow it. Then have them write the post-incident summary, and brief an AI stakeholder who wants to know when it will be fixed and whether it can happen again. The written note and the spoken briefing are not decoration — for most platform teams, the ability to explain an outage honestly to people who are not engineers is a substantial part of the job, and it is almost never assessed.
Why this work can be assessed
Incident work is a keyboard-and-conversation job — diagnose in a monitored sandbox, decide under incomplete information, then write the post-incident note and brief a stakeholder live. All three are directly reproducible.
Roles in this family
Sources
Every figure on this page is traceable. Where a claim could not be sourced it is stated qualitatively instead.
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Network and Computer Systems Administrators, 2025, https://www.bls.gov/ooh/computer-and-information-technology/network-and-computer-systems-administrators.htm
- Google Cloud, Announcing the 2025 DORA Report: State of AI-assisted Software Development, 24 September 2025, based on responses from nearly 5,000 technology professionals, https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report
- US Bureau of Labor Statistics, Occupational Outlook Handbook, Computer and Information Technology Occupations, 2025, https://www.bls.gov/ooh/computer-and-information-technology/home.htm
Hiring for one of these? We build the assessment for the specific role, run it under your brand, and return a ranked list with the evidence behind every score.
Book a walkthrough