Practical handbook
Performance Engineering and System Design Handbook
A senior-engineer field manual for designing and operating systems whose behavior is explainable, bounded, measurable, and economically justified.
Reading order
Contents
Part 1
Performance as a design discipline
8 entries
Part 1
Performance as a design discipline
- 01 The Performance Contract
- 02 Workload Models: Describe the Demand Before the System
- 03 Objectives, Indicators, and Budgets
- 04 Quantities, Units, and the Laws That Pay Rent
- 05 Distributions, Variance, and Tail Behavior
- 06 Queues, Utilization, and Backpressure
- 07 Bottlenecks, Constraints, and Causal Diagnosis
- 08 The Performance-Aware System Design Loop
Part 2
The machine beneath the system
8 entries
Part 2
The machine beneath the system
- 09 CPU Execution: From Instructions to Useful Work
- 10 Memory Hierarchy, Locality, and NUMA
- 11 Scheduling, Threads, and Concurrency Runtimes
- 12 Allocation, Garbage Collection, and Managed Runtimes
- 13 Storage Media and I/O Paths
- 14 Networks, Protocols, and the Cost of Distance
- 15 Virtualization, Containers, and Cloud Variability
- 16 Accelerators, Heterogeneous Compute, and Energy
Part 3
Reusable mechanisms and trade offs
13 entries
Part 3
Reusable mechanisms and trade offs
- 17 Data Layout, Algorithms, and Locality
- 18 Parallelism, Synchronization, and Contention
- 19 Asynchronous Execution, Queues, and Backpressure
- 20 Batching, Pipelining, Vectorization, and Amortization
- 21 Caching as a Consistency and Capacity Design
- 22 Serialization, Compression, and Protocol Shape
- 23 Load Balancing and Work Placement
- 24 Admission Control, Rate Limiting, and Load Shedding
- 25 Deadlines, Timeouts, Retries, Hedging, and Idempotency
- 26 Partitioning, Sharding, Skew, and Rebalancing
- 27 Replication, Quorums, and Read/Write Paths
- 28 Consistency, Coordination, and Transaction Boundaries
- 29 API, Schema, and Data-Contract Design for Performance
Part 4
Distributed architecture
9 entries
Part 4
Distributed architecture
- 30 Service Topology and Critical-Path Design
- 31 Messaging, Logs, and Delivery Semantics
- 32 Event Processing, Idempotent Effects, and ‘Exactly Once’
- 33 Time, Ordering, Identity, and Causality
- 34 Partial Failure, Tail Amplification, and Recovery
- 35 Geo-distribution and the Physics of Distance
- 36 Elasticity, Autoscaling, and Control Loops
- 37 Multi-tenancy, Isolation, and Fairness
- 38 State Placement, Ownership, and Derived Data
Part 5
System archetypes
10 entries
Part 5
System archetypes
- 39 Low-Latency Request/Response Services
- 40 Transactional Systems and OLTP
- 41 Key-Value Stores and Distributed Caches
- 42 Event Ingestion and Stream Processing
- 43 Batch, Analytical, and Data-Warehouse Systems
- 44 Search, Indexing, and Retrieval Systems
- 45 Object, File, Media, and CDN Systems
- 46 Real-Time Fan-out, Presence, and Collaboration
- 47 ML and AI Inference Systems
- 48 Edge, Mobile, and Intermittently Connected Systems
Part 6
Measurement modeling and validation
9 entries
Part 6
Measurement modeling and validation
- 49 Observability as a Performance Instrument
- 50 Profiling Across CPU, Memory, I/O, Locks, and Distributed Traces
- 51 Benchmarking and Microbenchmarking
- 52 Load, Stress, Spike, Soak, and Resilience Testing
- 53 Experimental Design and Performance Statistics
- 54 Capacity Planning and Demand Forecasting
- 55 Analytical Models and Queueing Networks
- 56 Simulation, Trace Replay, and What-If Analysis
- 57 Performance Regression Prevention and Continuous Validation
Part 7
Operation evolution and governance
8 entries
Part 7
Operation evolution and governance
- 58 Performance Incident Response
- 59 The Optimization Workflow: From Hypothesis to Durable Gain
- 60 Deployment, Rollout, and Online Migration
- 61 Resilience Under Overload and Recovery
- 62 Cost, Efficiency, and Sustainable Performance
- 63 Security, Privacy, and Performance Trade-offs
- 64 Performance Design Reviews and Decision Records
- 65 Ownership, Culture, and the Performance Program
Part 8
End to end case studies
9 entries
Part 8
End to end case studies
- 66 Case Study: Restoring a Latency SLO in a High-Throughput API
- 67 Case Study: Multi-Region Ordering and Payment
- 68 Case Study: The Hot-Key Cache Collapse
- 69 Case Study: A Burst-Tolerant Event Pipeline
- 70 Case Study: Deadline-Aware Search and Autocomplete
- 71 Case Study: Multi-Tenant Interactive Analytics
- 72 Case Study: Cost-Constrained AI Inference
- 73 Case Study: Zero-Downtime Storage and Schema Migration
- 74 Case Study: Anatomy of an Overload Incident
Reference section
Appendices
14 entries
Reference section
Appendices
- Appendix A Appendix A — Symbols, Units, and Magnitudes
- Appendix B Appendix B — Practical Probability and Statistics
- Appendix C Appendix C — Queueing and Capacity Formula Sheet
- Appendix D Appendix D — Orders of Magnitude
- Appendix E Appendix E — Performance Design Review Template
- Appendix F Appendix F — Benchmark and Experiment Report Template
- Appendix G Appendix G — Load-Test Plan Template
- Appendix H Appendix H — Performance Incident Runbook
- Appendix I Appendix I — Capacity-Planning Workbook Specification
- Appendix J Appendix J — Architecture Diagram Notation
- Appendix K Appendix K — Profiling and Diagnostic Command Cards
- Appendix L Appendix L — Glossary and Anti-Glossary
- Appendix M Appendix M — Review Question Bank and Design Drills
- Appendix N Appendix N — Further Reading Map