Books and videos I recommend, mostly on fundamentals. Some of these I've read cover to cover, others I only skimmed chapter by chapter. Either way, the most important part is putting it into practice. Otherwise you've wasted your time.
My criteria for picking a book: how much it improves my judgement per hour of reading, and how long that value lasts before the field moves on.
From 2007, and still true, because the hardware fundamentals haven't changed: caches, cache lines, TLBs, NUMA, prefetching. After reading it, performance work stops being guesswork. You start to see that data layout and access patterns are what actually matter.
If I could hand one book to every backend engineer, it would be this one. It covers replication, partitioning, transactions, consistency, and batch and stream processing, all explained through trade-offs instead of vendor marketing. It gives you the words to reason about any data system you'll run into.
Written by people from the Kotlin team, so it explains why the language works the way it does, not just the syntax. The fastest way I know to stop writing Java in Kotlin and start writing idiomatic Kotlin.
A reference, not a cover-to-cover read. I go back to it when I need to really understand a data structure or algorithm, instead of memorising patterns.
Most projects that fail, fail on what to build, not on how to build it. This book teaches you how to dig out, write down and validate requirements before anyone writes a line of code.
Read it after the one above: once you know what to build, this is how to model it. Ubiquitous language and bounded contexts alone are worth it, even if you skim the denser chapters.
A novel, so it's an easy read. It's about developer flow, feedback loops and what bureaucracy does to engineers. It helped me spot the same dysfunctions in real teams and gave me the words to talk about them.
About what happens to your code once it's in production. Circuit breakers, bulkheads, timeouts, and plenty of war stories. It makes you design for failure from day one.
How Google runs production: SLOs, error budgets, toil, on-call and blameless postmortems. Not every practice scales down to a small team, but the way of thinking does.
There are no best practices in architecture, only trade-offs. Practical advice on breaking up systems, splitting data, and choosing between orchestration and choreography in distributed architectures.
The companion to the SRE book. It treats security and reliability as one problem, because both are about how a system behaves under stress, and shows how to design for both from the start instead of bolting them on later.
How to build products on top of foundation models: evaluation, prompt engineering, RAG, agents, fine-tuning and inference optimisation. The best map I've found of a field that changes every month.
You implement a GPT-style model step by step in PyTorch. Once you've written attention and a training loop yourself, LLMs stop being magic.
Visual and practical. Clear illustrations of how tokens, embeddings and transformers work, followed by real uses like semantic search, classification and clustering.
Machine learning in production, beyond the model: data pipelines, features, deployment, monitoring and data drift. Most of the work is the system around the model, and this book is about that system.