Skip to content
Performance Engineering and System Design Handbook
Listening
Your library
Switch books
Performance Engineering and System Design Handbook
Listening now
Senior Engineering Interview Handbook
Listen
Solo Founder Product Engineering Handbook
Listen
Production Data Systems Handbook
Listen
Cybersecurity Engineering Handbook
Listen
AI Systems Handbook
Listen
The Rust Engineering Handbook
Open book
The Change Interface
Listen
CISSP Certification Guide
Open book
Project Management Mastery
Listen
Chapter
Preparing audio…
Read chapter
15s
30s
0:00
Choose Play when you are ready.
Text
Spoken text
Previous
Next
Playback speed
0.85×
1×
1.15×
1.25×
1.5×
2×
Sleep timer
Sleep
End of chapter
15 min
30 min
45 min
60 min
+5
Copy link
Volume
Spoken text is preparing…
Space
Play or pause
← →
Skip 15s / 30s
P N
Previous / next chapter
T
Spoken text
/
Search the text
M
Mute
0:00
/
0:00
Spoken text
Chapters
Follow spoken text
01
The Performance Contract
27:50
02
Workload Models: Describe the Demand Before the System
27:53
03
Objectives, Indicators, and Budgets
29:37
04
Quantities, Units, and the Laws That Pay Rent
31:08
05
Distributions, Variance, and Tail Behavior
31:25
06
Queues, Utilization, and Backpressure
28:35
07
Bottlenecks, Constraints, and Causal Diagnosis
32:16
08
The Performance-Aware System Design Loop
28:05
09
CPU Execution: From Instructions to Useful Work
20:29
10
Memory Hierarchy, Locality, and NUMA
30:00
11
Scheduling, Threads, and Concurrency Runtimes
32:00
12
Allocation, Garbage Collection, and Managed Runtimes
36:04
13
Storage Media and I/O Paths
35:21
14
Networks, Protocols, and the Cost of Distance
42:09
15
Virtualization, Containers, and Cloud Variability
40:34
16
Accelerators, Heterogeneous Compute, and Energy
42:27
17
Data Layout, Algorithms, and Locality
39:06
18
Parallelism, Synchronization, and Contention
42:37
19
Asynchronous Execution, Queues, and Backpressure
41:51
20
Batching, Pipelining, Vectorization, and Amortization
41:45
21
Caching as a Consistency and Capacity Design
53:59
22
Serialization, Compression, and Protocol Shape
52:21
23
Load Balancing and Work Placement
40:22
24
Admission Control, Rate Limiting, and Load Shedding
40:59
25
Deadlines, Timeouts, Retries, Hedging, and Idempotency
43:12
26
Partitioning, Sharding, Skew, and Rebalancing
43:55
27
Replication, Quorums, and Read/Write Paths
39:46
28
Consistency, Coordination, and Transaction Boundaries
44:09
29
API, Schema, and Data-Contract Design for Performance
43:41
30
Service Topology and Critical-Path Design
47:06
31
Messaging, Logs, and Delivery Semantics
35:05
32
Event Processing, Idempotent Effects, and ‘Exactly Once’
40:22
33
Time, Ordering, Identity, and Causality
44:30
34
Partial Failure, Tail Amplification, and Recovery
44:56
35
Geo-distribution and the Physics of Distance
44:40
36
Elasticity, Autoscaling, and Control Loops
46:29
37
Multi-tenancy, Isolation, and Fairness
40:18
38
State Placement, Ownership, and Derived Data
41:09
39
Low-Latency Request/Response Services
42:24
40
Transactional Systems and OLTP
45:22
41
Key-Value Stores and Distributed Caches
40:33
42
Event Ingestion and Stream Processing
42:36
43
Batch, Analytical, and Data-Warehouse Systems
32:51
44
Search, Indexing, and Retrieval Systems
34:44
45
Object, File, Media, and CDN Systems
32:13
46
Real-Time Fan-out, Presence, and Collaboration
31:32
47
ML and AI Inference Systems
32:50
48
Edge, Mobile, and Intermittently Connected Systems
37:04
49
Observability as a Performance Instrument
27:36
50
Profiling Across CPU, Memory, I/O, Locks, and Distributed Traces
32:49
51
Benchmarking and Microbenchmarking
30:50
52
Load, Stress, Spike, Soak, and Resilience Testing
39:17
53
Experimental Design and Performance Statistics
37:38
54
Capacity Planning and Demand Forecasting
42:43
55
Analytical Models and Queueing Networks
43:33
56
Simulation, Trace Replay, and What-If Analysis
44:11
57
Performance Regression Prevention and Continuous Validation
35:18
58
Performance Incident Response
45:53
59
The Optimization Workflow: From Hypothesis to Durable Gain
30:37
60
Deployment, Rollout, and Online Migration
31:17
61
Resilience Under Overload and Recovery
29:41
62
Cost, Efficiency, and Sustainable Performance
34:00
63
Security, Privacy, and Performance Trade-offs
28:31
64
Performance Design Reviews and Decision Records
30:47
65
Ownership, Culture, and the Performance Program
33:02
66
Case Study: Restoring a Latency SLO in a High-Throughput API
62:13
67
Case Study: Multi-Region Ordering and Payment
27:54
68
Case Study: The Hot-Key Cache Collapse
34:33
69
Case Study: A Burst-Tolerant Event Pipeline
55:22
70
Case Study: Deadline-Aware Search and Autocomplete
58:53
71
Case Study: Multi-Tenant Interactive Analytics
55:41
72
Case Study: Cost-Constrained AI Inference
60:23
73
Case Study: Zero-Downtime Storage and Schema Migration
42:22
74
Case Study: Anatomy of an Overload Incident
42:53