A Systems View of Efficient Machine Learning, from Silicon to Agents
Most of what is written about machine learning assumes the model is the interesting part and the machine is a detail. In practice the machine decides what you are allowed to build. This book is about the boundary where a model meets real hardware under a real budget for latency, memory, power and money: how to predict what that boundary will do to you, and what to change first when it does.
Fourteen parts, from the silicon up: rooflines, CPUs, GPUs and NPUs, kernels, compilers, quantisation, compression, vision, LLMs on small machines, robotics, profiling, serving and agents. 36 worked problems with every answer written out in full.
It's free. Start with Part 1, the roofline.
I write at usamah.me. A post to start with: Stop Latency Laundering.