Systems Foundations · core

Compute, Memory & Storage Hierarchy

CPU caches, NUMA, DRAM, GPU memory, NVMe, object storage, and the movement costs between them.

systems-foundationshardware

Mental model

Performance is usually data movement. Each level trades capacity and durability for latency and bandwidth, so placement decisions must follow the working set and access pattern.

How to study Compute, Memory & Storage Hierarchy

Begin by restating the mental model in your own words, then connect it to a concrete system you have built or operated. Name the mechanism, the constraint it addresses, and the trade-off it introduces. Use What Every Programmer Should Know About Memory to check details, but close the source before writing your explanation. Retrieval is the learning step; rereading is only preparation.

Next, compare Compute, Memory & Storage Hierarchy with Operating System Mechanics, Network Protocol Engineering. Ask what changes in correctness, latency, resource use, operability, and failure recovery. Complete Design exercise: Compute, Memory & Storage Hierarchy and preserve the command, input, output, and one failed attempt as evidence. Finish by explaining the idea without jargon to someone who has not studied the track.

Proof of understanding

  • Explain the mechanism from first principles and identify the state it reads or changes.
  • Give one situation where the concept is the right choice and one where it is not.
  • Predict a realistic failure mode before running the drill, then compare the prediction with evidence.
  • Connect the result to a roadmap or build artifact instead of treating the concept as isolated trivia.

Learn from primary sources

Practice and explain it back

Design exercise: Compute, Memory & Storage Hierarchy

CPU caches, NUMA, DRAM, GPU memory, NVMe, object storage, and the movement costs between them. Implement designOutline() returning non-empty values for: workingSet, dataMovement, bottleneck. Each value must name a concrete mechanism or decision.

Expected evidence: A design outline with workingSet, dataMovement, bottleneck plus an explicit failure mode or trade-off.

Open the interactive drill →

Review prompts

  • Two loops do the same arithmetic on the same array and one is many times slower. What is the usual cause?

Build evidence

Synthesize: Systems Foundations

Build a tiny HTTP/1.1 static-file server on raw TCP sockets without a framework or high-level HTTP server library. Parse requests, serve bounded files, handle partial I/O, inject failures, measure the result, and explain how the operating system, network, memory, concurrency, and storage paths interact.

  • Accepts TCP connections, parses a bounded HTTP GET request, serves fixture files, and returns explicit errors for malformed requests, missing files, and path traversal attempts
  • Names and implements a concurrency model with connection, request-size, timeout, and resource limits, including correct handling of partial reads and writes
  • Injects at least a slow client, malformed request, or interrupted transfer and demonstrates bounded failure and recovery
  • Reports a reproducible workload with throughput, p50/p95 latency, peak memory, and open-connection observations
  • Explains the loader, process, syscall, buffer, filesystem, TCP, and scheduling path in a concise architecture note

Trace a Tensor: Diagnose and Optimize One Workload

Trace one tensor-producing model operation from its numerical representation and computation graph through memory movement, kernel execution, engine scheduling, and request-level serving. Build or precisely model a reproducible workload, identify its dominant bottleneck, apply one justified optimization, and defend the resulting quality, latency, resource, and cost trade-offs.

  • Maps each lifecycle layer to the data representation, owner, work performed, and observable evidence
  • Runs or precisely models one reproducible workload and captures a before profile with latency and memory or bandwidth evidence
  • Diagnoses whether the dominant constraint is compute, memory movement, launch overhead, scheduling, or request shape
  • Applies one kernel, model-format, memory, or scheduling optimization and reports before/after measurements
  • Verifies output quality or numerical correctness and explains one trade-off, failure mode, or remaining risk

Prerequisites

Related concepts

Learning paths