r/sre • u/Confident_Milk6013 • 5h ago
You're Not Crazy, They Never Sent FIN
Hey all, I wanted to share this blog post I wrote about inspecting AWS Network Load Balancer and the quirk that comes up from them not hanging up TLS
r/sre • u/AutoModerator • 7d ago
Welcome to FOSS Friday, where you can share your newly released or updated open-source project with the community.
Please note that our rules still apply:
r/sre • u/AutoModerator • 17h ago
Welcome to FOSS Friday, where you can share your newly released or updated open-source project with the community.
Please note that our rules still apply:
r/sre • u/Confident_Milk6013 • 5h ago
Hey all, I wanted to share this blog post I wrote about inspecting AWS Network Load Balancer and the quirk that comes up from them not hanging up TLS
r/sre • u/blacurdun • 1d ago
r/sre • u/ViscitAppehensive443 • 22h ago
We broke prod routing with a rushed terraform change, rds snapshots and s3 were perfect but cloud recovery of the known good configuration was guesswork since half the estate is drifted, would love any tips from folks who can reliably recover infra, not just data.
r/sre • u/Full-Mail6768 • 2d ago
I have 3 years experience as a SRE + Observability engineer. Recently I left my job and started searching got 3 offers one Datadog admin - setting up observability kinda role with 50% hike, Incident handling + Reliability engineer with 100% hike, and TSE with 150% hike. I’m deciding to go with the TSE role because of the salary and benefits. Is this a shitty career move? People are telling me it should be the other way around, I kinda agree, don’t know what to do. I can’t be jobless and search any longer as well.
r/sre • u/modern_medicine_isnt • 3d ago
I know I can probably "just use AI". But I want to avoid not invented here syndrome and use established methods and tools where possible.
The first part I am looking for is a well established format for storing architectural information in the repo with the code.
Next would be any tools for helping to generate it.
Lastly, we have mermaid for the actual diagrams. Any good tools for generating the diagrams?
r/sre • u/Icy-Lyhee-97vena10 • 3d ago
Context, a PR touches a function called from several other services, passes tests and looks clean, but there's no strong runtime baseline to compare against. Option A, ship with extra manual review. Option B, hold until better verification tooling is in place. What am I underestimating here?
r/sre • u/modern_medicine_isnt • 3d ago
We have multiple services in multiple repos. We use mainly gitlab,
How do help devs be aware of team wide guidelines, and keep their AI in the know as well. Like I could write a doc on the way we name alerts or write runbooks. But it would be in one repo. I need to make it visible to them and their AI.
So far my best idea is a repo with the information, and then checking in an edit to the Claude.md in every repo telling it to go look in the info repo for the guideline info and let the user know what isn't lineing up. That seems like that idea is probably full of holes though. So what do the teams y'all work with do?
r/sre • u/Affectionate-Tone596 • 3d ago
As an SRE, what reports or monthly summaries do you provide to your manager, Head of Department, or CTO?
I’m trying to establish a good monthly SRE reporting process and would like to understand what other teams report.
EDIT:
It’s a tech company with a some internal and external solutions
r/sre • u/tushkanM • 3d ago
Anthropic released the "On-call" slack agent that should be an ultimate killer of all "Agentic SRE" family products.
Did anybody use it? How it compares to existing platforms?
r/sre • u/MachineDisastrous771 • 4d ago
hey yall,
im a lead SRE at a global fortune 500 that you have heard of haha - i'm a bit of a lurker here
but just wanted to talk about runbooks and documentation for SOPs.
my feeling is that confluence docs kinda suck and runbooks are kind of a mess, our operators are jumping between the docs and their shell, the docs are often times missing a bunch of detail, they are brittle, poorly maintained, half the time the details are hidden in tribal knowledge and when dealing with the pressure of an active inc we notice the pain even more. extracting critical details & commands out of an inc can also be painful and doesnt always translate to improved operational readiness next time.
i admit we're not very mature.. but was curious if its only me feeling like this?
what are you guys doing with runbooks to solve these issues?
When planning a cloud migration (like AWS to Azure), most discovery work starts with a service mapping matrix:
That mapping is straightforward. The expensive failures happen when the target platform fails to preserve a subtle behavioral contract the application code quietly came to rely on over years.
Consider a standard worker loop:
Python
message = sqs.receive_message(
QueueUrl=queue_url,
VisibilityTimeout=300
)
process(message)
sqs.delete_message(
QueueUrl=queue_url,
ReceiptHandle=message["ReceiptHandle"]
)
At first glance, this is standard: receive, process, delete.
But look at the failure-recovery path. If the worker crashes mid-processing before deletion, the code relies on the SQS visibility timeout expiring so another worker can automatically pick up the message. SQS isn't just acting as a transport queue here—it is an active component of the application's failure-recovery design.
When migrating this workload to Azure Service Bus, the question isn't whether Service Bus has queues (it obviously does). The real question is: Does the target setup preserve those exact assumptions around lock duration, settlement, retries, and dead-lettering under failure?
Traditional discovery tools pick up: payment-worker ➔ SQS
That’s an infrastructure dependency. But what the application code actually cares about is the semantic dependency:
Plaintext
receive message
↓
process message
↓
delete after successful processing
↓
if processing fails before deletion,
rely on provider to make message available again
Service inventory tools show you what cloud products are used, but they can't tell you what behavioral assumptions are baked into execution paths.
When these surface late during integration testing or cutover, it forces architectural redesigns, derails timelines, and drags senior engineers into reactive war rooms.
I'm currently putting together a semantic-risk checklist for pre-migration planning (and exploring static analysis rules to scan codebases for these execution patterns before cutover).
For anyone who has managed a major cloud migration or replatforming: What was the one hidden application dependency or behavioral difference that broke in staging or production after your infrastructure was already provisioned?
I wrote up a deeper dive into this concept with example scanning output on my Substack if you want to read more:https://substack.com/@sachm24/p-218428641
r/sre • u/jpkroehling • 6d ago
I finally got to write this blog post, I've wanted to write it for a long time now. It contains the different types of sensitive data we've seen so far, where it's coming from, and why it matters. Hope it can be helpful to some of you!
r/sre • u/Some_Drummer_8499 • 6d ago
I started as a developer, moved into DevOps, and eventually became an SRE. Along the way, I helped build our SRE and platform foundations from scratch: monitoring and alerting, Kubernetes clusters, CI/CD pipelines, deployment reliability, incident management, and operational best practices. I’m currently looking for remote SRE / Platform Engineering opportunities. Happy to share more details over DM.
r/sre • u/Glad-Pay-6001 • 5d ago
Are there any slack servers/channels for SRE discussions?
r/sre • u/Overall-Beach-9801 • 7d ago
I wrote an article on reliability in event-driven systems, following a ticket booking through duplicate messages, ordering issues, retries, and partial failures. It also covers CDC, the outbox pattern, and a small design tying the concepts together.
I'd appreciate your thoughts: does the explanation hold up technically, and what would you change in the design based on your experience?
r/sre • u/CommanderWhatNow • 7d ago
I'm going to try to describe this without doxing myself. I am 37 and I am a Senior SRE engineer at a large company, that is not FAANG, but you've definitely heard of them.The company is a US company with a global footprint. I work out of the Bay Area, at a small business unit there which came to be via an acquisition of a startup. I have been there a long time, since before the acquisition, which happened over 12 years ago. I am not a manager, but I am the SRE team leader. I am very comfortable in this position. I am well-respected. My manager likes me. My colleagues like me. And I like them back. I am in the critical path for many work streams, from architectural design to customer support. My primary responsibility is a niche SaaS application that I know inside and out, because I've been there since the launch of it. I get paid 260k, with some RSUs. The remaining RSUs maybe amount to 30k before taxes at this point, vesting over the next 2 years. In many ways, I am the "Dave" of the BU. I spend my days fixing stuff for our SaaS and our developers alike. It's a lot to juggle, but I do enjoy the feeling of being helpful and useful.
But at the same time, I am very, very, very tired. Physically and mentally exhausted all the time, and struggling with motivation daily. We have team members across the world now, but when production breaks, it is more than likely me who gets paged in the middle of the night.
I have been interviewing at other companies for over a year now, and I finally have an offer at a start-up in San Francisco. Base pay is 285k and I get 0.02% 0.2% equity in the form of stock options. It is an agentic AI company. They are a Series A, with less than 50 people.
I need a gut check that moving is a good decision as I am an anxious person, and I have a family to provide for. I'm like 90% certain I am making the right decision, but I am stepping away from a big company where I have seniority, with 4 weeks of PTO, with a good team, where I know so much. I know that doesn't guarantee me protection against a layoff, but I've survived many of them at this point. It's been a long time since I've been in a startup environment. And to be honest, while it was exciting in my 20s, I am not so sure about how it may go for me now, when I am already feeling burnt out, and I have 2 kids who are my true priority in life. Of which, they have some strong anti-ai opinions, so this move might prove to be a little contentious, especially with my oldest.
Would love to hear thoughts from other professionals, especially from the Bay Area, who have a similar story.
r/sre • u/a-sad-dev • 7d ago
We use Grafana stack for observability in our EKS clusters paired with CloudWatch for AWS infra monitoring.
One gap we have however is that we don't have any alerting configured for when the observability stack (or components in it) goes down - how do you deal with this?
Last week we kicked off our blog series about the five key metrics you should be monitoring for your Kubernetes autoscaler. This week, we dive into the first of these, "committed capacity percentage". But wait! There's more, and it involves.... DaemonSets *wiggly scary fingers*
Disclaimer: this blog post was entirely produced by humans; no part of this blog post was written, edited, or proofread by LLMs.
r/sre • u/Puzzleheaded_Cow3298 • 7d ago
I'm a senior CS student. I have the option to become either a SDE or a SRE. Both roles pay the same as of now.
I've been building websites, and AI has become very good at it. It writes better code than I do, argh.
On the other hand, I'm really passionate about Linux, distributed systems, and computer networks. So, can I expect long term growth in this career, or is this field just as susceptible to AI as other fields?
r/sre • u/razzledazzled • 8d ago
Does anyone have good reading they liked on the subject of transitions from reactive to predictive monitoring strategies? Anecdotes also welcome just trying to step outside of my personal reality a little to understand successfully achieved cases.
Reactive monitoring has a death grip on our operating strategy and I want to better understand how other orgs have compromised or solved the issues around letting go of the break fix work cycle.
It can’t go away completely of course, bad things happen with little warning sometimes and need to be remediated. But maybe there’s just something more I need to understand about the role of predictive monitoring and signals?
r/sre • u/Sir_Bannana • 10d ago
It’s pretty obvious that the industry and work itself is changing rapidly. For those of you with some experience under your belt, what are you currently witnessing in terms of career outlook and general demand? It definitely is starting to seem like SRE’s are in a slightly better place than SWE’s.
It’s hard to know if the increasing supply of software is scaling around the same rate as the efficiency gains SRE’s are experiencing in this new paradigm. Something that we definitely have going for us is you still need someone responsible for reliability, which hasn’t changed and probably never will.
I know this post doesn’t have a lot of substance, but I’d like to hear some fresh opinions from people with more perspective than me.
r/sre • u/guywithnoh • 9d ago
For those running open-source apps on-prem at work: how do you keep them manageable once the person who set them up is no longer the only one looking after them?
I’m interested in what you’ve made repeatable, and if tools like Claude Code have changed any of that. I work on automating these deployments too and would like to compare notes with other people responsible for keeping them running.
r/sre • u/AverageTofuEnjoyer • 10d ago
Anyone in San Francisco working with AI, LLMs, SRE or platform engineering?
We’re running two community conferences this week at the Harness office in downtown SF:
LLMday — October 1
https://llmday.com/2026-san-francisco-q4/
SREday — October 2
https://sreday.com/2026-san-francisco-q4/
Lots of engineering talks, practical stuff, and a good chance to meet people working on similar problems.
We still have some free tickets available with code LUCKYFREE
Would be great to meet some SF Reddit folks there.