
I’m Misha Strebkov - a Staff Software Engineer and technical lead at Google Cloud, where I work on the networking control plane and spend most of my time on what happens when large systems don’t behave.
That’s the thread through everything here. The happy path is the easy part; the engineering that matters lives in the failure modes - the regional outage, the network partition, the migration nobody adopts because it’s too painful. For the past few years I’ve been the technical lead for an organization-wide reliability program, started after a major infrastructure failure, working across hundreds of engineers and dozens of product teams to make systems survive the kind of fault that once disrupted an entire control plane.
A lot of my work is API design in service of reliability, and I care most about one principle: making the safe path the easy path - designing interfaces so the resilient choice is also the obvious, low-friction one, because that’s the only kind of safety that actually gets adopted. I’ve shipped APIs built on exactly that idea, now running in production at scale.
Before Google I was Director of Engineering at Jetlore (acquired by PayPal in 2018). I started out in Samara, Russia, where I grew up and earned my degree. Fourteen years in, my focus has settled on distributed systems, API design, reliability engineering, and large-scale control-plane architecture - mostly in Go and Java.
This is where I write about the parts of that work that rarely make it into a design doc: the trade-offs, the decision frameworks, the failures, and the unglamorous engineering that keeps big systems boring. If something here is useful, wrong, or worth arguing about, I’d like to hear it - I’m on LinkedIn.