Wix runs about 2.5% of the websites out there. When a site owner writes their own backend code, it runs on the Kubernetes platform I keep running, one site per gVisor sandbox. Dublin.
Untrusted code finds every crack. When a pod hangs and the cause isn't ours, I follow it down the stack until I find whose bug it is. Usually it's open source, so I fix it there and write up how I found it.
containerd once let nodes say Ready while no new pod could start. A snapshotter had hung and the snapshot GC waited on it forever. My fix bounded the GC, got merged, then got reverted because maintainers wanted it fixed on the client side. Fair. The deadline on proxy snapshotter calls went in instead.
The SOCI snapshotter leaked file descriptors on pod-start bursts until it hit its limit. One reader in file Verify never got closed. Now it does.
gVisor is where most of my time goes. A wedged sandbox could hold the shim lock forever, so runsc Kill, Stats and Status got 30s deadlines. A missed systrap interrupt could block teardown, so now only the stuck subprocess dies, not the sandbox. The last indefinite waits are still in review.
metrics-server drops the whole node from kubectl top after one slow kubelet scrape. Open.
The Security Profiles Operator is where it started, smaller: AppArmorProfile owners in node status, 2023.
My Claude Code mod. It draws replies in color right in the terminal: tables, code with copy buttons, mermaid diagrams and charts, 15 themes.
It started because I'd built a whole VS Code wrapper just to see mermaid diagrams in Claude Code replies. A mod turned out to be the better way.