Seven For Sunday: August 2019
Updated:
On Retroing
This week’s outage at Cloudflare took down a significant percentage of the internet. Cloudflare has posted a placeholder post-mortem that hopefully will continue to flesh out details of what went wrong and why.
I think there’s a lot to learn in studying specific instances of fragility. At my shop we host blameless post-incident retros that utilize the 5 Whys technique for each outage we have. While nobody hopes for incident, I look forward to our retros to learn, to find out where and why I've been wrong, to measure the maturity of my shop’s ability to retro and learn, and to spawn conversations with my fellow technologists about fragility and ways to harden or mitigate risk from it.
Cloudflare seems likely to IPO soon. I look forward to owning shares.
Summer 2019 Reading List
Five things I've read & enjoyed lately:
1) Morgan Housel at Collaborative Fund had a post this week called Why Things Break. Several layers of abstraction away from software engineering, nonetheless I think there’s a lot to learn in studying generalized fragility.
This post enumerates five ways success can lead to fragility:
- Success pushes you away from whatever made you successful to begin with
- Success increases size, size increases complexity, complexity plants landmines
- Success teaches you how to win the last war, which becomes the only war you know how to fight
- Success reduces the impression of needing room for error
- What looked like success was random, mismeasured, or a temporary trend
2) Crucial Conversations, both the book and the need to have them, came up at work this week.
If you haven’t read the book, the authors define a crucial conversation as
A discussion between two or more people where the stakes are high, opinions vary, and emotions run strong.
I’ve got this theory: Healthy shops are organizationally architected for the minimum number of crucial conversations. If you could plot the number of Crucial Conversations that need to happen in your shop over time, you’d have an inverse graph of your organizational health.
WARNING: There is a consulting industry that’s emerged around this book, that, if you Google, you will quickly encounter. I’m neutral-to-wary.
3) On Marcus and Friends by Vicki Boykis
I’m a firm believer in finding someones who are doing what doing so much better than you’re doing it, and Vicki’s Normcore Tech is an inspiration for where I hope to take this blog. Just this one post has a full week’s worth of linky goodness.
4) Shane Parrish’s podcast interview with author Jim Dethmer contained a framework for stratifying motivation that I’m pondering how to use in my life and work.
Sources of Motivation (in order from lowest to highest) per Jim Dethmer:
- Fear, guilt, shame, anger, and rage
- Extrinsic rewards (fame, money)
- Intrinsic rewards (doing things for the sake of a purpose or calling) (the line between healthy and unhealthy motivations is here)
- Play (when your work feels like a child at play)
- Love (love of the thing)
I appreciate how Podcast Notes helps me make sense of what I've listened to and my own notes on a podcast.
5) Zac J. Szewczyk writes - on his own self-run website because meta - Run Your Own Website.
This caught my eye on Hacker News. Timely for me, considering what I’m doing here.
Spending time considering whether an anchor tag belongs inside or outside of a headline tag reminds me of a zen admonition.
Product Management In 2019
I’ve been considering Product Management a lot since my current shop instantiated a Product team two years ago. Revisiting canonical definitions helps ground both my expectations and my sense for what my partners in building software expect from me.
My own general framework for re-grounding in a concept is
- Wikipedia article
- One to three authoritative, high SEO posts… including comments and/or tweets about them
- The most authoritative book on the subject
And then it’s a question of looking for consistencies and shared understanding (or lack thereof) across a community to assess maturity of concept and start mining useful frameworks.
Thoughts On Burnout
This is part of a broader theme I’m picking up on re: the worrying state of mental health in the developed world.
I appreciated this post because it including multiple perspectives, including some that cast light on a "risk of medicalizing everyday distress… labelling more and more healthy people as sick and building bigger potential markets for those selling medicines"
Here’s what I know about burnout:
My dad worked for the federal government. QA for the DOD. He made sure the people who made the missiles and tanks were making them safe, performant, reliable. That just sounds like the type of job that would require weekend slack or phone call sessions. I never once saw my dad work on the weekends. Not a phone call, not a whoops Dad has to go into the office for a half-day Saturday. He did his 9-5 (actually 7-3), came home and relaxed with his family.
Most of my career has been spent in small-shop scenarios: Solopreneur or Founder or CTO. There’s a 24/7 responsibility there, even if it’s rarely exercised. I once spent an afternoon and evening in an Internet Cafe in Vienna diagnosing and fixing a database problem, vacation on pause. Losing a day off seemed like a reasonable price for the rewards running my own shop provided.
I’ve done this in six different roles. Eight years being the longest stretch. By year seven I was Burned. The. Fuck. Out. I had to make a change, so I did: I booked a six month sabbatical. (Aside: This turned into a month sabbatical, five months getting a new venture I was excited about off the ground, and a lesson for myself in how I will struggle if I ever try to retire)
Today I work at a mid-sized org that’s part of a very large enterprise. Redundancy exists… nothing I do can’t be done by anyone else. I hold very few if any tribal secrets that nobody else knows. Even if my entire team were to disappear, we write enough documentation and develop apps and services in such common patterns that someone else could eventually be found to debug almost any problem.
And yet, I find myself having to remind myself to show some discipline. To not check Slack on a sunny Sunday morning, or Saturday night after coming home from the bar.
I try to be aware of myself as model. If I’m doing the _inspire_ part of my job, others are going to emulate my behavior. Nobody can tell when I’m monitoring a weekend incident response or catching up on interesting posts, but at least by modeling not posting, I can avoid setting the example that hey, Engineering Leaders here work on weekends and thereby maybe, just maybe, make a tiny contribution to alleviating burnout in the industry.
Local Portland Org I'm Excited About
MilkRun is an online marketplace and grocery delivery service that connects consumers and Chefs directly with hundreds of local farmers. They’re located here in PDX. I love ideas around using tech to bring people closer to the producers of goods.
Reading
The Ultimate Guide to Structuring a 90-Day Onboarding Plan by Alida Miranda-Wolff.
Onboarding is something I’ve spent a lot of time this year thinking about. We want to invest in onboarding suffices at the Strategic level, leaving much of the burden of So how the hell do we do that to the Operations level: largely to Engineering Managers and their Tech Leads.
This week I wrote a guide for my fellow Engineering Managers on how to onboard at my shop. Not a lot novel; just getting people in, helping them meet greet get tools get training get learning and get going. Alida Miranda-Wolff's post was a primary source for my thinking.
Conference I (Almost) Attended
OSCON, the O’Reillly Open Source conference that once again took place here in Portland.
My company sponsors and speaks regularly at OSCON. This year we had individual contributors speak on a project they’ve open sourced. Open source contributions are a nascent thing for us, and having engineers speak about what they've contributed represents, to me, a significant incremental improvement in our community participation. I didn’t get to attend, but I did host a dinner for some of the out of town attendees at one of my favorite restaurants in Portland or anywhere.
Large-scale Incident I'm Learning From
We’ve had a period of fragility at work, nothing out of the ordinary for a fast-diverging tech shop but still something we’ve had strategic focus on and operational imperatives towards solving. So I’ve been consuming content about how large-scale failures are handled.
The Challenger disaster and the subsequent Rogers Commission Report are a large-scale example of incident and post-incident response. The story continues to the Columbia disaster and reveals how little of the Rogers Commission Report recommendations were implemented, but what was interesting to me here is it got me thinking about a few things: Who plays the role of Richard Feynman in my shop? Does every incident need a Feynman, or are some or even many incidents straightforward enough to not require it? What is the cost of someone going full-on Feynman, not just in time but on morale? Etc.
Analogy I'm Considering
The importance of pruning roses.
One of my co-workers used the analogy of pruning roses too describe the importance of stopping initiatives, even ones that might seem to be succeeding, if in doing so we would make either the plant healthier or encourage proliferation of new growth.
When I was teaching creative writing, Kill Your Babies / Kill Your Darlings was my favorite unit. Assertive and at times aggressive editing is one of the shortest paths to improving your writing. The resistance to deleting words you’ve worked hard to write seems perfectly Freudian and related to a fear what if the faucet won’t turn on tomorrow, what if I can’t write words this well again. At the org level, we see similar hesitation to kill projects and initiatives that have an acquired momentum.
Terminology Refinement
Monitoring and Observability by Cindy Sridharan.
Cindy Sridharan is on my “must follow on twitter and medium” shortlist, and this week she referenced a post she’d written that made me laugh, because we’ve been going through the exercise of operationally orienting around foundational engineering 'imperatives' this year, and transforming our language from Monitoring to Observability has been part of that.
The goals of “monitoring” and “Observability” are different. “Observability” isn’t a substitute for “monitoring” nor does it obviate the need for “monitoring”; they are complementary. “Observability” might be a fancy new term on the horizon, but it really isn’t a novel idea. Events, tracing, exception tracking are all a derivative of logs, and if one has been using any of these tools, one already has some form of “Observability”.
Thoughts On Performance And Perception In Software
Fast Software, the Best Software by Craig Mod
A few selected quotes include:
"Fastness in software is like great margins in a book — makes you smile without necessarily knowing why.”
"...slowness feels indicative of unseen rot on the inside of the machine...”
"Sublime Text has - in my experience - only gotten faster. I love software that does this: Software that unbloats over time.”
"A typewriter is an excellent tool because, even though it’s slow in a relative sense, every aspect of the machine itself operates as quickly as the user can move… The best software inches ever closer to the physical directness of something like a typewriter.”
Funny enough, while I was working on this, Chrome froze on my MacBook. I was on a flight and probably should have expected the browser, untethered from the internet, might choke with eight tabs full of hungry, connectivity-presumed apps open. Nonetheless… it struck me as unfaithful act.
Process
Bets, Not Backlogs from Shapeup by Ryan Singer
One of my teams is drowning in process. I've been trying to gather as much context as I can about different approaches to the process of making software. Health, happiness and productivity of my teams are first-class concerns when managing engineers - a lot of my note taking while reading posts like this involves how a particular approach might impact them.
Interior Design
Psychology Of Colors In A Workplace by Jeffrey Fermin
Sometimes I need a read that reminds me of simple design facets and their effects on us. This is that.
I’m wondering what this means for those of us who for extra-large orgs where not using on-brand colors is not an option.
Short But Sweet
Yes, No, Maybe by Hugh MacLeod
The Value Of Readable Code
Easier Error Handling Using Async/Await by Jesse Warden
One of my leads has been hinting that I need to spend some time with Node’s asyc/await and some of the changes coming in version 12 and as 8 reaches EOL, and I finally did so this week.
As a JavaScript programmer and aficionado, I’d be somewhat disappointed if this change results in a Go-like exceptionless paradigm. I’ve long considered over-handling exceptions an anti-pattern, so changes that encourage it have my antennae up.
Book I'm Reading (Part 1)
The Clock Of The Long Now: Time and Responsibility by Stewart Brand
Book I'm Reading (Part 2)
The Book of Life: Daily Meditations with Krishnamurti by Jiddu Krishnamurti
These two books are dancing together in my mind. Long Now challenges the reader to consider a twenty-thousand year span 'now'. Book of Life challenges the reader to listen and be present in every moment. Read one way, they might seem to operate in mutually exclusive zones. Looked at another way, there's a very interesting and fun overlap developing in my mind.