Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

>If the root cause is to be fixed someone needs to look at it in depth

The root cause has already been looked at in depth 999 times when the same issue has come up. It's already been RCAed and the fix has been put in the backlog to be implemented sometime next year. In the meantime while we wait for the fix, we will continue to do a full, ad-hoc RCA every time the exact same issue appears, with the exact same results every time, because managers genuinely think it is a valuable way to spend our time.

I understand your point, but the relative utopia of a team you're describing is not really the situation I'm talking about. We have on-call periods where the exact same issue will appear 10-20 times per week, and each and every time it is treated as a completely novel issue with an ad-hoc response, even though we already know beforehand what the root cause is and what the fix is. It's an incredible waste of time and contributes significantly to on-call engineers being overloaded, and yet we continue to do it and then are baffled when all of our engineers leave the team due to being overworked.

There's also nothing excluding runbooks and root cause analyses from existing together, either. In fact, most good runbooks specifically include steps to determine when an RCA is necessary and how to conduct one. There really is no excuse to not use runbooks as much as possible. If over-reliance on runbooks is having a negative impact due to engineers not applying personal judgement, then that is certainly an issue to be addressed, but the answer is almost never to completely abolish runbooks and documentation.



> the exact same issue will appear 10-20 times per week, and each and every time it is treated as a completely novel issue with an ad-hoc response

Yeah, this sounds like a very bad situation where management won't let you do something that reduces ops pain because it isn't the most desirable solution, but they won't let you prioritize the right solution either. The next thing that happens is that on-call folks develop ad-hoc quasi-runbooks and share them amongst a subset of people (or just keep them to themselves to make their own life easier) and those quasi-runbooks become critical to ops, but not documented or shared by everyone. It's pure dysfunction.


This does sound pretty dysfunctional. You'd think that for something that's causing 999 on-calls getting the root cause fixed would be a priority. What I described obviously falls apart when the team has no ability to actually fix issues. Perhaps the original intent was to get those issues fixed but that somehow got lost as the org grows larger.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: