Skip to content

Practical handbook

Performance Engineering and System Design Handbook

A senior-engineer field manual for designing and operating systems whose behavior is explainable, bounded, measurable, and economically justified.

A layered performance observatory tracing workload through queues, machine resources, distributed paths, telemetry, and bounded feedback loops.

Reading order

Contents

Part 1

Performance as a design discipline

8 entries
  1. 01 The Performance Contract
  2. 02 Workload Models: Describe the Demand Before the System
  3. 03 Objectives, Indicators, and Budgets
  4. 04 Quantities, Units, and the Laws That Pay Rent
  5. 05 Distributions, Variance, and Tail Behavior
  6. 06 Queues, Utilization, and Backpressure
  7. 07 Bottlenecks, Constraints, and Causal Diagnosis
  8. 08 The Performance-Aware System Design Loop

Part 2

The machine beneath the system

8 entries
  1. 09 CPU Execution: From Instructions to Useful Work
  2. 10 Memory Hierarchy, Locality, and NUMA
  3. 11 Scheduling, Threads, and Concurrency Runtimes
  4. 12 Allocation, Garbage Collection, and Managed Runtimes
  5. 13 Storage Media and I/O Paths
  6. 14 Networks, Protocols, and the Cost of Distance
  7. 15 Virtualization, Containers, and Cloud Variability
  8. 16 Accelerators, Heterogeneous Compute, and Energy

Part 3

Reusable mechanisms and trade offs

13 entries
  1. 17 Data Layout, Algorithms, and Locality
  2. 18 Parallelism, Synchronization, and Contention
  3. 19 Asynchronous Execution, Queues, and Backpressure
  4. 20 Batching, Pipelining, Vectorization, and Amortization
  5. 21 Caching as a Consistency and Capacity Design
  6. 22 Serialization, Compression, and Protocol Shape
  7. 23 Load Balancing and Work Placement
  8. 24 Admission Control, Rate Limiting, and Load Shedding
  9. 25 Deadlines, Timeouts, Retries, Hedging, and Idempotency
  10. 26 Partitioning, Sharding, Skew, and Rebalancing
  11. 27 Replication, Quorums, and Read/Write Paths
  12. 28 Consistency, Coordination, and Transaction Boundaries
  13. 29 API, Schema, and Data-Contract Design for Performance

Part 4

Distributed architecture

9 entries
  1. 30 Service Topology and Critical-Path Design
  2. 31 Messaging, Logs, and Delivery Semantics
  3. 32 Event Processing, Idempotent Effects, and ‘Exactly Once’
  4. 33 Time, Ordering, Identity, and Causality
  5. 34 Partial Failure, Tail Amplification, and Recovery
  6. 35 Geo-distribution and the Physics of Distance
  7. 36 Elasticity, Autoscaling, and Control Loops
  8. 37 Multi-tenancy, Isolation, and Fairness
  9. 38 State Placement, Ownership, and Derived Data

Part 5

System archetypes

10 entries
  1. 39 Low-Latency Request/Response Services
  2. 40 Transactional Systems and OLTP
  3. 41 Key-Value Stores and Distributed Caches
  4. 42 Event Ingestion and Stream Processing
  5. 43 Batch, Analytical, and Data-Warehouse Systems
  6. 44 Search, Indexing, and Retrieval Systems
  7. 45 Object, File, Media, and CDN Systems
  8. 46 Real-Time Fan-out, Presence, and Collaboration
  9. 47 ML and AI Inference Systems
  10. 48 Edge, Mobile, and Intermittently Connected Systems

Part 6

Measurement modeling and validation

9 entries
  1. 49 Observability as a Performance Instrument
  2. 50 Profiling Across CPU, Memory, I/O, Locks, and Distributed Traces
  3. 51 Benchmarking and Microbenchmarking
  4. 52 Load, Stress, Spike, Soak, and Resilience Testing
  5. 53 Experimental Design and Performance Statistics
  6. 54 Capacity Planning and Demand Forecasting
  7. 55 Analytical Models and Queueing Networks
  8. 56 Simulation, Trace Replay, and What-If Analysis
  9. 57 Performance Regression Prevention and Continuous Validation

Part 7

Operation evolution and governance

8 entries
  1. 58 Performance Incident Response
  2. 59 The Optimization Workflow: From Hypothesis to Durable Gain
  3. 60 Deployment, Rollout, and Online Migration
  4. 61 Resilience Under Overload and Recovery
  5. 62 Cost, Efficiency, and Sustainable Performance
  6. 63 Security, Privacy, and Performance Trade-offs
  7. 64 Performance Design Reviews and Decision Records
  8. 65 Ownership, Culture, and the Performance Program

Part 8

End to end case studies

9 entries
  1. 66 Case Study: Restoring a Latency SLO in a High-Throughput API
  2. 67 Case Study: Multi-Region Ordering and Payment
  3. 68 Case Study: The Hot-Key Cache Collapse
  4. 69 Case Study: A Burst-Tolerant Event Pipeline
  5. 70 Case Study: Deadline-Aware Search and Autocomplete
  6. 71 Case Study: Multi-Tenant Interactive Analytics
  7. 72 Case Study: Cost-Constrained AI Inference
  8. 73 Case Study: Zero-Downtime Storage and Schema Migration
  9. 74 Case Study: Anatomy of an Overload Incident

Reference section

Appendices

14 entries
  1. Appendix A Appendix A — Symbols, Units, and Magnitudes
  2. Appendix B Appendix B — Practical Probability and Statistics
  3. Appendix C Appendix C — Queueing and Capacity Formula Sheet
  4. Appendix D Appendix D — Orders of Magnitude
  5. Appendix E Appendix E — Performance Design Review Template
  6. Appendix F Appendix F — Benchmark and Experiment Report Template
  7. Appendix G Appendix G — Load-Test Plan Template
  8. Appendix H Appendix H — Performance Incident Runbook
  9. Appendix I Appendix I — Capacity-Planning Workbook Specification
  10. Appendix J Appendix J — Architecture Diagram Notation
  11. Appendix K Appendix K — Profiling and Diagnostic Command Cards
  12. Appendix L Appendix L — Glossary and Anti-Glossary
  13. Appendix M Appendix M — Review Question Bank and Design Drills
  14. Appendix N Appendix N — Further Reading Map