262 questions across 20 topics - answered to the depth an interviewer actually expects.
Role tracks: Platform Engineer (junior β staff) Β· DevEx Β· Platform Architect Β· Kubernetes Platform Β· Platform Security Β· Platform SRE Β· AWS Β· Azure Β· GCP Β· FinOps Β· Platform Lead
Every answer gives you a short answer you can say out loud, the detail and trade-offs behind it, a runnable example, and the follow-ups to expect.
Pick your role Β· Browse topics Β· All questions Β· Knowledge graph Β· How answers are structured Β· Contributing
π Explore the knowledge graph β - all 262 questions as a live map of the concepts they share.
β Star the project if it helps you land the role.
Platform engineering interviews are not infrastructure interviews with a new title. They probe a specific set of judgements: whether you can design an interface other engineers will voluntarily use, where you put a guardrail versus a golden path, how you version a contract forty teams depend on, and how you know whether any of it worked.
This guide is organised around those judgements. It assumes you already know what a Pod is, and spends its pages on the decisions above that line - tenancy boundaries, control-plane design, escapable abstractions, rollout policy, cost attribution, and the operating model of a team whose customers are other engineers.
New to the field, or moving in from DevOps or SRE? Start with Platform Engineering Fundamentals, then Developer Experience. The vocabulary in those two topics is assumed by everything else.
Eleven tracks, each a reading order rather than a pile of links. Start at the left and work right.
Interview in a fortnight? Start with Platform Engineering Interviews - it covers the loop structure, the platform design round, how to present work you have done, and what to ask back, cross-linked to the rest of the guide.
Grouped by theme, with question counts and difficulty mix. Click a topic to open its index, which opens with what interviewers probe there.
262 questions across 20 topics - π’ 100 Beginner Β· π‘ 82 Intermediate Β· π΄ 80 Advanced
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| Platform Engineering Fundamentals | 13 | 5 | 5 | 3 | what a platform actually is, why self-service is the defining property, and where platform⦠|
| Developer Experience | 13 | 5 | 6 | 2 | cognitive load, service catalogues, portals, scorecards, onboarding time, and the metrics that⦠|
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| Platform Architecture | 13 | 5 | 4 | 4 | control plane versus data plane, the developer-facing API, escapable abstractions, interface⦠|
| Multi-Tenancy and Isolation | 13 | 5 | 4 | 4 | tenancy models, namespace versus cluster boundaries, noisy neighbours, quotas, network⦠|
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| Kubernetes Platform | 13 | 5 | 3 | 5 | cluster topology, node pools and scheduling, fleet upgrades, admission control as a platform⦠|
| Control Planes and Abstractions | 13 | 5 | 3 | 5 | Crossplane, custom resources as platform contracts, workload specifications, Terraform at scale,β¦ |
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| GitOps and Continuous Delivery | 13 | 5 | 5 | 3 | Argo CD and Flux, repository structure at scale, environment promotion, secrets, drift⦠|
| Progressive Delivery and Feature Flags | 14 | 5 | 3 | 6 | flags as a platform capability, flag debt and its SLAs, rollout curves with abort criteria,β¦ |
| Environments and Ephemeral Infrastructure | 13 | 5 | 5 | 3 | preview environments per pull request, data seeding without leaking production, environment⦠|
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| Platform Security | 14 | 5 | 3 | 6 | workload identity without long-lived keys, secretless pipelines, secrets management,β¦ |
| Policy as Code and Governance | 13 | 5 | 5 | 3 | OPA and Kyverno, rolling out policy without breaking teams, compliance frameworks mapped to⦠|
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| AWS Platform Engineering | 13 | 5 | 4 | 4 | multi-account organisation design, account vending, EKS pod identity, runtime selection,β¦ |
| Azure Platform Engineering | 13 | 5 | 4 | 4 | management groups and landing zones, Azure Policy guardrails, AKS workload identity, runtime⦠|
| GCP Platform Engineering | 13 | 5 | 4 | 4 | the folder and project hierarchy, organisation policies, Workload Identity Federation, GKE⦠|
| Multi-Cloud and Hybrid Platforms | 13 | 5 | 4 | 4 | genuine drivers versus slogans, the cost of portable abstractions, cross-provider primitive⦠|
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| Platform Reliability | 13 | 5 | 3 | 5 | SLOs for a platform rather than an app, error budget policy, platform on-call, incidents where⦠|
| Platform Observability | 13 | 5 | 4 | 4 | OpenTelemetry collectors as shared infrastructure, cardinality and cost control, zero-effort⦠|
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| Platform FinOps | 13 | 5 | 4 | 4 | cost attribution for shared infrastructure, showback versus chargeback, unit economics,β¦ |
| Platform Team and Operating Model | 13 | 5 | 4 | 4 | Team Topologies, staffing, adoption metrics, migration without a mandate, deprecation, the RFC⦠|
| Topic | Questions | π’ | π‘ | π΄ | What it covers |
|---|---|---|---|---|---|
| Platform Engineering Interviews | 13 | 5 | 5 | 3 | the loop structure, the platform design round, presenting your own platform work,β¦ |
Every question in the repository, collapsed by topic - open only the ones you are studying.
26 questions
Platform Engineering Fundamentals Β· 13 questions Β· π’ 5 π‘ 5 π΄ 3
Open the Platform Engineering Fundamentals index β
Developer Experience Β· 13 questions Β· π’ 5 π‘ 6 π΄ 2
Open the Developer Experience index β
26 questions
Platform Architecture Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the Platform Architecture index β
Multi-Tenancy and Isolation Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the Multi-Tenancy and Isolation index β
| No. | Question | Difficulty |
|---|---|---|
| 40 | What is multi-tenancy and why do platforms need it? | π’ Beginner |
| 41 | What is a tenant, and how do you define one in a platform? | π’ Beginner |
| 42 | What does a Kubernetes namespace isolate, and what does it not? | π’ Beginner |
| 43 | What are resource quotas and limit ranges? | π’ Beginner |
| 44 | How does role-based access control scope what a tenant can do? | π’ Beginner |
| 45 | What tenancy models can a platform offer? | π‘ Intermediate |
| 46 | How do you isolate tenants at the network layer? | π‘ Intermediate |
| 47 | How do virtual clusters change the tenancy trade-off? | π‘ Intermediate |
| 48 | How do Pod Security Standards harden a shared cluster? | π‘ Intermediate |
| 49 | When do you give a tenant its own cluster instead of a namespace? | π΄ Advanced |
| 50 | How do you stop one tenant degrading another? | π΄ Advanced |
| 51 | How do you isolate tenant identity and data? | π΄ Advanced |
| 52 | How do you design tenant onboarding and offboarding? | π΄ Advanced |
26 questions
Kubernetes Platform Β· 13 questions Β· π’ 5 π‘ 3 π΄ 5
Open the Kubernetes Platform index β
Control Planes and Abstractions Β· 13 questions Β· π’ 5 π‘ 3 π΄ 5
Open the Control Planes and Abstractions index β
| No. | Question | Difficulty |
|---|---|---|
| 66 | What is infrastructure as code? | π’ Beginner |
| 67 | What is a reconciliation loop? | π’ Beginner |
| 68 | What is a custom resource definition? | π’ Beginner |
| 69 | What is the difference between Terraform and OpenTofu? | π’ Beginner |
| 70 | What is Terraform state and why does it need protecting? | π’ Beginner |
| 71 | What problem does a workload specification like Score solve? | π‘ Intermediate |
| 72 | What changed in Crossplane v2 and why does it matter? | π‘ Intermediate |
| 73 | How do kro, Crossplane compositions, and Helm charts differ for composing resources? | π‘ Intermediate |
| 74 | What is Crossplane and how does it differ from Terraform? | π΄ Advanced |
| 75 | How do you model a platform API with Kubernetes custom resources? | π΄ Advanced |
| 76 | How do you manage Terraform at platform scale? | π΄ Advanced |
| 77 | How do you handle resource deletion safely in a control plane? | π΄ Advanced |
| 78 | How do you provide self-service infrastructure without handing out cloud credentials? | π΄ Advanced |
40 questions
GitOps and Continuous Delivery Β· 13 questions Β· π’ 5 π‘ 5 π΄ 3
Open the GitOps and Continuous Delivery index β
Progressive Delivery and Feature Flags Β· 14 questions Β· π’ 5 π‘ 3 π΄ 6
Open the Progressive Delivery and Feature Flags index β
| No. | Question | Difficulty |
|---|---|---|
| 92 | What is progressive delivery? | π’ Beginner |
| 93 | What is a feature flag? | π’ Beginner |
| 94 | What is a canary deployment? | π’ Beginner |
| 95 | What is a blue-green deployment? | π’ Beginner |
| 96 | What is the difference between deploying and releasing? | π’ Beginner |
| 97 | When do you use a feature flag instead of a canary deployment? | π‘ Intermediate |
| 98 | How do you test code that sits behind feature flags? | π‘ Intermediate |
| 99 | What is OpenFeature and why does a vendor-neutral flag API matter? | π‘ Intermediate |
| 100 | How do you run feature flags as a platform capability? | π΄ Advanced |
| 101 | How do you manage feature flag debt? | π΄ Advanced |
| 102 | How do you design a progressive rollout and its abort criteria? | π΄ Advanced |
| 103 | How do you design a kill switch you can trust? | π΄ Advanced |
| 104 | How do you keep percentage rollouts consistent across services? | π΄ Advanced |
| 105 | What can a feature flag not roll back? | π΄ Advanced |
Environments and Ephemeral Infrastructure Β· 13 questions Β· π’ 5 π‘ 5 π΄ 3
Open the Environments and Ephemeral Infrastructure index β
27 questions
Platform Security Β· 14 questions Β· π’ 5 π‘ 3 π΄ 6
Open the Platform Security index β
Policy as Code and Governance Β· 13 questions Β· π’ 5 π‘ 5 π΄ 3
Open the Policy as Code and Governance index β
52 questions
AWS Platform Engineering Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the AWS Platform Engineering index β
Azure Platform Engineering Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the Azure Platform Engineering index β
GCP Platform Engineering Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the GCP Platform Engineering index β
Multi-Cloud and Hybrid Platforms Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the Multi-Cloud and Hybrid Platforms index β
26 questions
Platform Reliability Β· 13 questions Β· π’ 5 π‘ 3 π΄ 5
Open the Platform Reliability index β
| No. | Question | Difficulty |
|---|---|---|
| 198 | What are SLIs, SLOs, and SLAs? | π’ Beginner |
| 199 | What is an error budget? | π’ Beginner |
| 200 | What is a blameless postmortem? | π’ Beginner |
| 201 | What are RTO and RPO, and how do they differ from high availability? | π’ Beginner |
| 202 | What is toil and how does a platform team reduce it? | π’ Beginner |
| 203 | How do you run on-call for a platform team? | π‘ Intermediate |
| 204 | How do you use chaos engineering to test platform guarantees? | π‘ Intermediate |
| 205 | How do you run a production readiness review? | π‘ Intermediate |
| 206 | How do you define SLOs for a platform rather than an application? | π΄ Advanced |
| 207 | What does an error budget policy look like for a platform team? | π΄ Advanced |
| 208 | How do you run an incident when the platform itself is the incident? | π΄ Advanced |
| 209 | How do you make a platform degrade gracefully instead of failing closed? | π΄ Advanced |
| 210 | How do you plan disaster recovery for the platform itself? | π΄ Advanced |
Platform Observability Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the Platform Observability index β
26 questions
Platform FinOps Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the Platform FinOps index β
| No. | Question | Difficulty |
|---|---|---|
| 224 | What is FinOps and what part of it does a platform team own? | π’ Beginner |
| 225 | What are the main cloud pricing models? | π’ Beginner |
| 226 | What is cost allocation tagging and why does it fail? | π’ Beginner |
| 227 | What is rightsizing? | π’ Beginner |
| 228 | Why is Kubernetes cost allocation hard? | π’ Beginner |
| 229 | What is the difference between showback and chargeback, and which works? | π‘ Intermediate |
| 230 | How do you find and remove idle platform cost? | π‘ Intermediate |
| 231 | What is the FOCUS specification and why does it matter? | π‘ Intermediate |
| 232 | How do you put cost feedback into the developer workflow? | π‘ Intermediate |
| 233 | How do you attribute shared platform cost to teams? | π΄ Advanced |
| 234 | How do you build unit economics for a platform? | π΄ Advanced |
| 235 | How do you use commitments and spot capacity on behalf of every team? | π΄ Advanced |
| 236 | How do you manage the cost of GPU and AI workloads on a platform? | π΄ Advanced |
Platform Team and Operating Model Β· 13 questions Β· π’ 5 π‘ 4 π΄ 4
Open the Platform Team and Operating Model index β
| No. | Question | Difficulty |
|---|---|---|
| 237 | What does Team Topologies say about platform teams? | π’ Beginner |
| 238 | What does a platform product manager do? | π’ Beginner |
| 239 | What is the difference between a platform team and an infrastructure team? | π’ Beginner |
| 240 | What is an enabling team and how does it work with a platform team? | π’ Beginner |
| 241 | How do you gather feedback from platform users? | π’ Beginner |
| 242 | How do you size and staff a platform team? | π‘ Intermediate |
| 243 | How do you measure platform adoption and success? | π‘ Intermediate |
| 244 | How do you run an RFC process for platform decisions? | π‘ Intermediate |
| 245 | How do you fund a platform team? | π‘ Intermediate |
| 246 | How do you migrate teams onto the platform without a mandate? | π΄ Advanced |
| 247 | How do you deprecate a platform capability? | π΄ Advanced |
| 248 | How do you build a platform roadmap and say no? | π΄ Advanced |
| 249 | How do you scale a platform organisation from one team to several? | π΄ Advanced |
13 questions
Platform Engineering Interviews Β· 13 questions Β· π’ 5 π‘ 5 π΄ 3
Open the Platform Engineering Interviews index β
Every answer follows the same four beats, so you can read one section deep or all four:
| Section | What it gives you |
|---|---|
| Short answer | Two or three sentences you could say out loud. Start here. |
| Detail | The substance - mechanisms, trade-offs, and the vocabulary that signals experience. |
| Example | Real manifests, policies, commands, or diagrams you can run and adapt. |
| Interview tips | The follow-up questions, the common traps, and the points that separate a strong answer from a memorised one. |
Difficulty is marked π’ Beginner Β· π‘ Intermediate Β· π΄ Advanced. Every topic carries at least five Beginner, three Intermediate, and two Advanced questions, so each one can be read as a path: the Beginner answers build the vocabulary, and the Intermediate and Advanced answers are about judgement rather than recall - which is where most platform roles, filled at senior and above, are actually decided.
Three ways to work through it:
- Preparing for a specific role - take the track from Pick your role, and read each topic README's "What interviewers probe here" before its questions.
- Broad revision - read the short answers across a topic, then go deep only where you hesitate.
- Obsidian / note vault - every file carries YAML frontmatter (
title,id,category,difficulty,tags), so the whole repository can be dropped into a vault and browsed by tag.
Platform engineering questions do not sit in neat boxes: tenancy decides your isolation boundary, which decides your policy model, which decides what your golden path can promise. The knowledge graph makes those dependencies visible - every question is a node, and two questions are linked when they genuinely argue about the same concept.
The links are derived, not hand-maintained. Each answer is scanned for the ~70 concepts this field turns on (control plane, golden path, policy as code, blast radius, error budget, chargeback, β¦), and a pair is linked when the concepts they share are rare across the vault. Rarity weighting is what stops "kubernetes" linking everything to everything; length normalisation stops the longest answers becoming hubs.
python3 scripts/build_knowledge_graph.py --output docs/index.html # the deployed page
python3 scripts/build_knowledge_graph.py --stats # edge counts, hubs, orphans
python3 scripts/build_knowledge_graph.py --json docs/graph.json # the raw graph
python3 scripts/inject_wikilinks.py # write the same edges into the question filesThree renderings ship with the builder and are selectable with --prototype:
| Prototype | What it is for |
|---|---|
constellation |
(deployed) Force-directed map of the whole vault, clustered by theme. Click a question to light up its neighbourhood. |
atlas |
The capability map: groups and topics laid out in reading order, with concept links drawn on demand. |
pathway |
Difficulty lanes that generate a study route - the questions to read before the hard one makes sense. |
The page is a single self-contained HTML file with no dependencies, deployed to GitHub Pages by knowledge-graph.yml on every push to main.
inject_wikilinks.py writes the same edges into each question's ## Related Questions block, between <!-- RELATED:START --> and <!-- RELATED:END --> markers, as both [[wikilinks]] (for Obsidian-style vaults) and relative links (for GitHub). It is idempotent, and --check fails CI when a block has drifted from the graph.
.
βββ platform-engineering-fundamentals/ # topic-slug/
β βββ README.md # generated topic index
β βββ what-is-platform-engineering.md # question-slug.md
βββ ...
βββ scripts/
β βββ lib_content.py # shared frontmatter/vault parsing
β βββ generate_indexes.py # regenerates all indexes from question files
β βββ validate_content.py # CI validation of the whole vault
β βββ build_knowledge_graph.py # derives the concept graph, renders the HTML site
β βββ inject_wikilinks.py # writes graph edges into `## Related Questions` blocks
β βββ topic_meta.json # topic registry: order, group, description, study notes
βββ docs/
β βββ index.html # generated knowledge graph, deployed to GitHub Pages
βββ .github/workflows/
βββ validate-and-format.yml # runs validation + Prettier on every PR
βββ knowledge-graph.yml # rebuilds and deploys the graph on push to main
Directories and filenames carry no numeric prefixes - they are pure slugs. Ordering comes from two places instead:
- Topic order - the
orderfield inscripts/topic_meta.json, which is also the registry of which directories count as topics. - Question order - the
idfield in each question's frontmatter, unique across the repository.
Renaming or reordering is therefore a metadata edit, not a mass file rename.
Every question file starts with frontmatter:
---
title: "What is a golden path and how does it differ from a mandate?"
id: 4
category: "Platform Engineering Fundamentals"
difficulty: "Intermediate"
tags:
- platform-engineering
- platform-engineering-fundamentals
- interview-questions
---The indexes are generated, not hand-written. The question files are the single source of truth; topic READMEs and the tables above are rendered from their frontmatter. After adding or editing a question:
python3 scripts/generate_indexes.py # rewrite all indexes
python3 scripts/inject_wikilinks.py # refresh the related-questions blocks
python3 scripts/build_knowledge_graph.py --output docs/index.html # rebuild the graph page
python3 scripts/validate_content.py # verify frontmatter, naming, links, index freshnessBoth are stdlib-only Python 3.11+ - no dependencies to install. CI runs the same commands and fails the pull request on drift.
New questions, better answers, and corrections are all welcome. CONTRIBUTING.md covers the file format, naming rules, and the local checks to run before opening a pull request.
Not sure where to start? Open an issue - there are templates for a new question, a correction to an existing answer, a new topic, and broken links or tooling.
Two documents set the ground rules: the Code of Conduct - including the rules on interviewer privacy, NDAs, and plagiarism - and the Security Policy, which covers how to report a committed credential, a workflow vulnerability, or a dangerously over-permissive example policy privately.
Three sibling repositories, same structure, same tooling - pick the one that matches the role you are interviewing for:
- Ultimate DevOps Guide - DevOps, SRE, DevSecOps, and cloud engineering. Start there for container, CI/CD, Linux, and networking fundamentals; this guide assumes them.
- Ultimate AI Engineering Guide - LLM fundamentals, prompt engineering, RAG, agents and MCP, fine-tuning, evaluation, and LLMOps. The counterpart for AI platform and AI engineering loops, where this guide covers the infrastructure the models run on.
- Ultimate Platform Engineering Guide (you are here) - the platform layer between the two: golden paths, control planes, tenancy, and the operating model of a team whose customers are other engineers.
The vocabulary this guide uses is not invented here. It leans on the work of the platform engineering community:
- Team Topologies by Matthew Skelton and Manuel Pais - platform teams as enabling teams, cognitive load as the design constraint, and the thinnest viable platform.
- CNCF Platforms Working Group - the platform maturity model and the capability-plane vocabulary.
- Google SRE - SLOs, error budgets, and the toil framing that platform reliability builds on.
- OpenFeature, OpenTelemetry, Crossplane, Backstage, and Argo CD - the open standards and projects most of these answers reference.
- DORA - the delivery metrics used throughout to argue about outcomes rather than output.
Released under the MIT License.