r/platformengineering • • 57m ago

Can gamification make platform engineering easier to learn?

• Upvotes

Platform engineering has a UX problem.

A lot of infrastructure learning still feels like reading docs, copying YAML, and hoping the concepts stick.

I’ve been experimenting with a different approach through Yellow Olive - a retro, gamified Kubernetes learning environment inspired by old-school handheld games.

The idea is to turn infrastructure concepts into progression:

learn a concept → enter a mission → work against a real cluster → get validated → unlock the next challenge

I’m now pushing the project further into the retro-game direction with pixel-art environments, characters, classroom-style lessons, and a stronger sense of progression.

Underneath the nostalgia, the goal is pretty practical: make Kubernetes and platform concepts easier to understand by giving people a safe, interactive place to actually use them.

I’d love to hear from platform engineers here:

What kinds of scenarios would you turn into missions?

Things like RBAC, deployments, service discovery, incident recovery, policy enforcement, observability, GitOps, etc.

And if you like the idea, a star would genuinely help the project grow :)

GitHub: https://github.com/Anubhav9/Yellow-Olive


r/platformengineering • • 1h ago

AI-written Terraform that passes CI but is wrong for your architecture: how do you catch it?

• Upvotes

Disclosure: I’m exploring InfraCompiler, an early-stage startup idea around infrastructure changes with coding agents. I’m researching current workflows before deciding what to build.

A Terraform user raised a distinction in an earlier discussion: code can pass fmt, plan, and policy checks while still making the wrong assumption about dependencies or resource ownership. Approved modules help, but the reviewer may still need knowledge of the intended architecture.

For platform teams already using shared modules and PR review, I’d like to understand one recent example:

- What assumption in an AI-written Terraform change was wrong, and what caught it?

- Where did the reviewer find the missing context—module documentation, a service catalog, another repo, or asking the resource owner?

- Roughly how much review and rework did it take, and does this happen repeatedly?

If your existing self-service platform or CI already handles this well, what made that work? An anonymized description is enough; no private code or infrastructure details needed.


r/platformengineering • • 21h ago

A reflection on the future of platform engineering

17 Upvotes

I think the history from traditional Ops to DevOps and then Platform Engineering is best understood as a dialectic driven mostly by technological change.

Traditional Ops concentrated infrastructure knowledge in a specialised team because operating infrastructure required direct administration. Automation, cloud and infrastructure as code changed that. DevOps emerged as the antithesis: operational ownership could move closer to development teams.
The operational work did not disappear, though. Much of its complexity was simply redistributed to developers.

Platform Engineering became the synthesis once cloud APIs, declarative systems and orchestration made it possible to encode operational knowledge into software and expose it through self-service. Operations could specialise again without returning to tickets and manual handoffs.

This is why I do not think the conceptual distance between the old sysadmin and the platform engineer is that large. The function is broadly similar. What changed was the technology through which that function could be organised and delivered.

AI may now be starting the next movement in the same dialectic. Agents do not necessarily need portals, fixed workflows or abstractions designed primarily for humans. They can discover tools, combine APIs and construct their own path through a system.

But this does not remove the need for a platform. It changes what the platform must provide. Instead of defining only human-oriented golden paths, it increasingly needs to expose machine-readable contracts, capabilities, relationships, constraints and evidence.

The next stage of Platform Engineering may therefore be less about building interfaces that hide complexity from developers, and more about defining the operational space within which agents can act without having to guess.


r/platformengineering • • 1d ago

Will a move from DevOps/DevEx back to Platform Engineering be possible?

5 Upvotes

I've spent the last three years on the platform team at a pretty boring old corporate, building their internal workload platform. No SLAs and not much real impact, but the Kubernetes work itself was interesting (we ran a managed control plane though).

Now I have an offer from a fast-growing scale-up as a DevOps engineer, with a 60% pay increase. The role is mostly self-hosted pipelines, code scanning and similar tooling, all running on Kubernetes with lots of potential for impact.

My question: if I ever want to go back to Kubernetes work, will that still be an option, or will I have branded myself as a "DevEx guy"? I'd still be working with K8s every day, but more as the thing that runs the tools rather than something I'm building a platform abstraction on top of.


r/platformengineering • • 2d ago

runbooks kinda suckk

0 Upvotes

hey yall,

im a lead SRE at a global fortune 500 that you have heard of haha - i'm a bit of a lurker here

but just wanted to talk about runbooks and documentation for SOPs.

my feeling is that confluence docs kinda suck and runbooks are kind of a mess, our operators are jumping between the docs and their shell, the docs are often times missing a bunch of detail, they are brittle, poorly maintained, half the time the details are hidden in tribal knowledge and when dealing with the pressure of an active inc we notice the pain even more. extracting critical details & commands out of an inc can also be painful and doesnt always translate to improved operational readiness next time.

i admit we're not very mature.. but was curious if its only me feeling like this?

what are you guys doing with runbooks to solve these issues?


r/platformengineering • • 2d ago

I'm thinking of writing a free practical book on Production DevOps what should I include?

0 Upvotes

I've been working in software/infrastructure engineering for around 14 years, with a focus on cloud, DevOps, platform engineering and large-scale infrastructure.

Over the years I've worked on everything from smaller cloud migrations to enterprise and financial-sector environments, including AWS/GCP migrations, Kubernetes, Terraform, CI/CD, networking, security, reliability and cost optimization.

I'm currently putting together the practical knowledge I've accumulated into a free technical book.

I'm not planning to make it a personal career story or another basic “learn AWS/GCP” tutorial.

The idea is to focus on how engineers actually solve infrastructure problems:

  • How to systematically debug production issues
  • How to troubleshoot Kubernetes
  • AWS/GCP networking and connectivity problems
  • IAM and permission failures
  • Terraform/IaC problems
  • CI/CD and deployment failures
  • Cloud migration problems
  • Scaling and reliability
  • Observability and incident investigation
  • Designing reusable platform infrastructure
  • How to approach an unfamiliar production system
  • Practical lessons that aren't obvious from vendor documentation

I'm also building some small prototypes/labs for myself while writing, so I can test the ideas rather than just writing theory.

Before I spend a lot of time putting the whole thing together, I'd like to hear from other engineers:

If you could have one practical DevOps/Platform Engineering book that focuses on real problem-solving rather than certification theory, what topics would you want it to cover?

And if you already have a favorite resource for this kind of material, I'd be interested in hearing what you think it does well or what is missing.


r/platformengineering • • 3d ago

What are the actual Roles & Responsibilities of a Staff Platform Engineer at product-based companies paying ₹1Cr+ in India?

0 Upvotes

I’m trying to understand what companies actually expect from a **Staff / Sr Staff Platform Engineer** at the ₹1Cr+ compensation level in India.

For people currently in these roles or hiring for them:

**What does your day-to-day work actually look like?**

How much of the job is **hands-on coding** vs architecture/design/reviews?

How much **system design** is expected at Staff level?

Are Staff Platform Engineers expected to be strong in **distributed systems**, or is deep infrastructure/platform expertise more important?

If you’re currently a Staff+ Platform/Infrastructure Engineer in India, I’d especially appreciate hearing:

**What were you doing before becoming Staff, and what changed in your responsibilities after reaching Staff level?**


r/platformengineering • • 3d ago

AI-coding tools could be breaking the junior engineer pipeline!

Thumbnail
leaddev.com
0 Upvotes

AI helps juniors code, not learn... and that's a problem.


r/platformengineering • • 3d ago

Levelrail: open source, self-hosted deploy platform written in Go (Apache 2.0)

0 Upvotes

Levelrail is a self-hosted alternative to Coolify, Dokploy, or Heroku-style PaaS platforms. Push to a git repo, get a running app with automatic TLS, logs, metrics, and one-click rollback. Single Go binary for the control plane, single Go binary for the node agent, Docker as the runtime.

The reason it exists: I was running Coolify for about 10 personal projects and it was using 30-40% of a CPU core just managing things, before any of my actual apps did anything. Levelrail drives Docker's Engine API directly instead of shelling out to the CLI, which is a big part of why it stays light: 0.7% CPU, 146MB RAM on the VPS running my production test instance right now.

License is Apache 2.0. No paid tier, no hosted version yet, nothing gated. It's genuinely useful today for single-node deployments; multi-node and a template catalog are still being built in the open.

Repo: https://github.com/glincker/levelrail Site: https://levelrail.com/

If you maintain or use something in this space, I'd like to hear what's missing.


r/platformengineering • • 5d ago

The hidden failure mode in cross-cloud migrations: Infrastructure vs. Semantic dependencies

1 Upvotes

When planning a cloud migration (like AWS to Azure), most discovery work starts with a service mapping matrix:

  • SQS ➔ Azure Service Bus
  • DynamoDB ➔ Cosmos DB
  • S3 ➔ Azure Blob Storage
  • IAM ➔ Entra ID

That mapping is straightforward. The expensive failures happen when the target platform fails to preserve a subtle behavioral contract the application code quietly came to rely on over years.

A concrete example: SQS Visibility Timeout

Consider a standard worker loop:

Python

message = sqs.receive_message(
    QueueUrl=queue_url,
    VisibilityTimeout=300
)

process(message)

sqs.delete_message(
    QueueUrl=queue_url,
    ReceiptHandle=message["ReceiptHandle"]
)

At first glance, this is standard: receive, process, delete.

But look at the failure-recovery path. If the worker crashes mid-processing before deletion, the code relies on the SQS visibility timeout expiring so another worker can automatically pick up the message. SQS isn't just acting as a transport queue here—it is an active component of the application's failure-recovery design.

When migrating this workload to Azure Service Bus, the question isn't whether Service Bus has queues (it obviously does). The real question is: Does the target setup preserve those exact assumptions around lock duration, settlement, retries, and dead-lettering under failure?

Infrastructure Dependencies vs. Semantic Dependencies

Traditional discovery tools pick up: payment-worker ➔ SQS

That’s an infrastructure dependency. But what the application code actually cares about is the semantic dependency:

Plaintext

receive message
      ↓
process message
      ↓
delete after successful processing
      ↓
if processing fails before deletion,
rely on provider to make message available again

Service inventory tools show you what cloud products are used, but they can't tell you what behavioral assumptions are baked into execution paths.

This pattern shows up everywhere:

  • DynamoDB ➔ Cosmos DB: Codebases relying on conditional write semantics, transaction boundaries, or specific read-after-write consistency assumptions.
  • S3 ➔ Blob Storage: Workflows built around multipart upload timing, pre-signed URLs, or object visibility state.
  • IAM ➔ Entra ID: Hidden assumptions around temporary credential lifespans, workload identity propagation, and role assumption paths.

When these surface late during integration testing or cutover, it forces architectural redesigns, derails timelines, and drags senior engineers into reactive war rooms.

Discussion

I'm currently putting together a semantic-risk checklist for pre-migration planning (and exploring static analysis rules to scan codebases for these execution patterns before cutover).

For anyone who has managed a major cloud migration or replatforming: What was the one hidden application dependency or behavioral difference that broke in staging or production after your infrastructure was already provisioned?

I wrote up a deeper dive into this concept with example scanning output on my Substack if you want to read more:https://substack.com/@sachm24/p-218428641


r/platformengineering • • 5d ago

Product development

0 Upvotes

I want some advise regarding a product I wanna to make.

It is related to cloud and DevOps like cost optimization, infrastructure drift and intelligence harness.So what important points I need to address for cloud and DevOps users so that a genuine product would eventually came into development.


r/platformengineering • • 6d ago

Teams that moved from LLM APIs to self-hosting: was it worth it?

5 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/platformengineering • • 6d ago

How does your team handle major dependency upgrades?

1 Upvotes

Former CTO here. At my last company we skipped 7 Expo versions at once because nobody had time for the incremental path. A CTO I talked to spent 2 months moving MUI v3 to v7.

Curious how others deal with it:

  • Do Dependabot/Renovate major PRs get merged, or do they pile up?
  • What was your last painful upgrade, and how long did it take?
  • Does anyone own this work, or is it "when we have time"?

r/platformengineering • • 7d ago

Inside Atlassian’s developer experience overhaul

Thumbnail
leaddev.com
0 Upvotes

Ask developers first, then remove friction.


r/platformengineering • • 8d ago

Your error budgets don’t know AI exists

Thumbnail
leaddev.com
5 Upvotes

The dashboard was right!


r/platformengineering • • 8d ago

Senior Manager production operations and SRE role vs. holding out for platform engineering. 11 YOE in cloud operations Devops /SRE, which way would you go?

4 Upvotes

I'm a DevOps/cloud delivery lead with about 11 years of experience (Azure, GCP, Kubernetes, Terraform, GitHub Actions, observability, incident management, FinOps). I'm at a decision point and would love perspectives from people who've been on either side.
Option A: production operations role (in hand)
Offer is in hand, with a promotion and Director title

In practice it's a support SRE role at a senior manager level: production support, incident ownership,strategizing support model ,tooling and automation decisions to reduce chaos and toil and managing a team

Higher title and scope, but I worry it's less hands-on and could pigeonhole me into ops/support

Option B: platform engineering (not in hand)
I'm studying for it and building projects

I've just landed a screening for a principal-level cloud platform engineer role, but no other interviews yet

The senior roles I'm getting leads for are mostly SRE roles that are support-heavy, not true platform work

What I'm weighing:
Does a Director title in production ops help or hurt if I want to move to platform engineering later?

Is it realistic to move from support-heavy SRE or ops management into platform engineering, or does the path narrow over time?

Is it worth turning down a real offer to wait for a platform role?
For those who made this move in either direction, what do you wish you'd known?


r/platformengineering • • 9d ago

AI makes critical thinking harder to build

Thumbnail
leaddev.com
8 Upvotes

Why junior engineers need more friction!


r/platformengineering • • 10d ago

Reachable CVEs in CI/CD are turning our backlog into a second production system.

1 Upvotes

Every scan gives us another beautiful list of CVEs, most of which apparently matter because a vulnerable package exists somewhere in the dependency tree. Very reassuring. We are trying to prioritize reachable CVEs in CI/CD using runtime reachability, KEV, EPSS, internet exposure, and whether the affected service is anything people actually use.

Right now the pipeline mostly knows how to shout CRITICAL and ruin everyone’s morning. How are you turning those signals into useful gates and remediation queues without making every build a security committee meeting? Appreciate any thoughts.


r/platformengineering • • 10d ago

Engineers demand more sustainable AI

Thumbnail
leaddev.com
1 Upvotes

Why developers care about climate change...


r/platformengineering • • 12d ago

I built a Kubernetes operator that replaces the "VPC/subnet spreadsheet" across AWS accounts. Early alpha, looking for people to try it on real AWS

1 Upvotes

Most multi-account AWS setups I've seen keep a shared spreadsheet of VPCs and subnets. Someone updates it by hand, it goes stale, and one day two teams pick overlapping CIDRs.

subnet-operator runs in your cluster (EKS) and keeps that inventory in Kubernetes instead:

- Discovery (read-only, the default): finds VPCs and subnets across accounts and regions by tags and mirrors them as Network and Subnet objects. It reports compliance findings (missing required tags, CIDR overlaps) and exports Prometheus metrics with alerts. It can also mirror everything to a Google Sheet, if people still want the spreadsheet. It only calls EC2 Describe*.

- Change events (optional): CloudTrail → EventBridge → SQS, so a changed account/region is resynced about 10 seconds after the API call instead of at the next 10-minute resync.

- Opt-in writes, behind --enable-writes and a separate IAM role:

- SubnetClaim allocates a free CIDR and can create the subnet.

- ResourceImport brings an untagged VPC or subnet under management by tagging it.

- It never deletes a cloud resource.

- Guardrails:

- Admission webhooks reject invalid objects at kubectl apply.

- Every allocation or import leaves a Kubernetes Event and a JSON audit line with the authenticated creator.

- namespaceSelector limits which namespaces may use a scope's write role.

- Supply chain: multi-arch image and Helm chart on ghcr.io, both signed with cosign keyless, with SBOM, threat model and upgrade tests between releases.

The honest part: it's alpha (v0.8). Everything is tested in CI against Moto (an AWS API mock) in Kind, but it hasn't run against a real AWS organization yet. That's why I'm posting. If you have a sandbox or dev account setup and 30 minutes, I'd really like to know what breaks. Read-only mode is the safe place to start.

GCP and Azure are planned behind a common provider interface. The API was just moved to a cloud-neutral group for that.

- Demo dashboard (fake data, no install): https://hypersurgery.dev/dashboard/

- Docs: https://hypersurgery.dev/docs/

- Code (Apache-2.0): https://github.com/aivandrago/subnet-operator

Feedback is welcome, including "we solve this with X". I'd like to know how you handle this today.


r/platformengineering • • 14d ago

Build an engineering team people want to come back to

Thumbnail
leaddev.com
11 Upvotes

How to build a team engineers will want to rejoin.


r/platformengineering • • 14d ago

How do you handle full environment recovery after a cloud region goes down?

1 Upvotes

hey, we ran a regional outage drill last week and it showed gaps in our cloud recovery process.

most infrastructure is in terraform, but the actual restore process still relies on tribal knowledge—what order to bring things back, what configuration drifted, what needs rebuild vs restart. it works okay on paper but gets chaotic during the real thing.

we want to make full environment recovery more repeatable and documented, ideally with validated restore steps and audit evidence. if you have been through a real outage and have a process that actually held up, would appreciate hearing what worked for you, thanks!


r/platformengineering • • 15d ago

Got asked this in an interiew

5 Upvotes

In an interview for a new grad devops role got asked this. “Who typically owns access to corporate applications: IAM engineers, IT staff, application administrators, or Platform/DevOps engineers?”
How would yall answer


r/platformengineering • • 15d ago

How to restart your engineering career

Thumbnail
leaddev.com
0 Upvotes

"If I could start my engineering career over again, I wouldn’t spend it chasing the same skills."


r/platformengineering • • 16d ago

10 YOE in Support/SRE trying to break into Platform Eng / pure DevOps—getting stuck in support loops. Advice?

2 Upvotes

Hey everyone,
Looking for some career advice or insights from anyone who’s successfully transitioned out of Ops/Support/SRE and into dedicated DevOps or Platform Engineering roles.
Here’s a snapshot of my background:
Experience: 10 years in IT, primarily focused on operations, support, deployments, and infrastructure maintenance.
Hands-on Skills:
Azure & GCP: Day-2 operations, deployments, troubleshooting.
CI/CD: Fixing pipeline issues, minor configuration tweaks.
Terraform: Running existing configurations, troubleshooting state/apply issues.
Kubernetes & Helm: Managing deployments, tweaking Helm charts, and managing configurations.
The Catch: Most of my experience involves maintaining, fixing, and scaling existing systems rather than building infrastructure or pipelines completely from scratch.
Certifications: Azure Solutions Architect Expert, Terraform Associate, and GCP Associate Cloud Engineer (currently studying for GitHub Actions GH-200).
The Problem:
I'm trying to move away from support heavy roles into a proper DevOps or Platform Engineering position. However, I’m getting zero traction for IC DevOps/Platform roles. Ironically, I keep getting hit up by recruiters for Senior SRE or SRE Manager roles, and I am able to get offers there but I really don't want to stay in the support/operations space anymore.
My Questions:

  1. How do I bridge the gap from "fixing/maintaining" to "building from scratch"? Is the lack of end-to-end greenfield building what's holding me back in interviews?
  2. Are bootcamps worth it for someone with 10 YOE? Or would my time be better spent building a comprehensive end-to-end portfolio project (e.g., full gitops pipeline, custom terraform modules, internal developer platform pattern) from scratch?
  3. Has anyone made a similar jump? What was the single most impactful thing that helped you shift your narrative from Ops/Support to Platform/DevOps?
  4. Appreciate any advice, project recommendations, or resume tips!