Pathubs AI Engineering Curriculum โ€ข Phase 01: Foundations

Computer Fundamentals: Hardware Architecture, Memory Systems & Accelerators

Master modern computer architecture from first principles. Understand how the physical computer executes code under the hood: the difference between low-latency CPUs, massively parallel GPUs, and low-power NPUs; the memory hierarchy from registers to NVMe SSDs; memory capacity versus memory bandwidth; PCIe data-movement bottlenecks; digital representations from bits to floating-point tensors; and real-world hardware bottleneck diagnostics.

โฑ๏ธ Estimated Time:50 Minutes
๐ŸŽฏ Level:Foundational to Intermediate
๐Ÿ“Š Track:AI Engineering & Systems Foundations
โœจ Mode:Interactive Workload & Architecture Labs
01

The Computer Mental Model & Modern Execution Layers

How physical silicon components coordinate to bring software instructions and AI models to life.

At its most fundamental level, every computing device follows the classic Input โ†’ Processing โ†’ Memory โ†’ Storage โ†’ Output paradigm. An input device captures data (keystrokes, audio, or network packets), the processing units perform mathematical operations, high-speed memory buffers active instructions, non-volatile storage preserves state across power cycles, and output interfaces deliver results.

However, modern AI engineering requires looking past this simplistic abstraction. An AI engineer does not interact directly with raw silicon; software runs through an orchestrated hardware-software stack:

Modern AI System Execution Hierarchy
๐Ÿ Layer 5: Application & Frameworks
Python Runtime, PyTorch, Hugging Face Transformers, vLLM, Triton
๐Ÿ’ป Layer 4: Operating System & Device Drivers
Linux / Windows Kernel, Virtual Memory Manager, NVIDIA CUDA Driver, DirectML
โšก Layer 3: Hardware Accelerators
CPU (Latency-optimized cores), GPU (Tensor/Matrix cores), NPU (Neural ASIC)
๐Ÿ”‹ Layer 2: Memory Subsystem
CPU Caches (L1/L2/L3), Host DDR5 RAM, High-Bandwidth Accelerator VRAM (GDDR6X, HBM3e)
๐Ÿ’พ Layer 1: Storage & Interconnects
PCIe Gen 4/5 Bus, NVMe M.2 SSDs, Motherboard Chipset, Power Delivery (VRMs/PSU)

Key Physical Hardware Roles for AI Engineers

Hardware ComponentPhysical RoleWhy It Matters for AI Engineering
CPU (Central Processor)Executes general-purpose operating system logic and sequential application instructions.Orchestrates data loaders, tokenizes text, coordinates Python execution, and dispatches compute kernels to accelerators.
GPU (Graphics Processor)Massively parallel compute array containing thousands of arithmetic logic units (ALUs) and Tensor Cores.Executes matrix multiplications ($W \times X$) that power neural network training and high-throughput batch inference.
NPU (Neural Processor)Specialized low-power application-specific integrated circuit (ASIC) designed for matrix tensor math.Delivers 40โ€“60+ TOPS of continuous inference (speech recognition, computer vision) at 5โ€“15W without draining laptop batteries.
System RAM (DDR5)High-speed volatile memory directly connected to the CPU memory controller.Holds training datasets, Python objects, operating system buffers, and models during initial disk loading.
GPU VRAM (GDDR6X / HBM3e)Ultra-high bandwidth memory attached directly to the graphics processor silicon die.Houses active model weights, KV caches, and intermediate activation tensors during execution.
NVMe SSD StorageNon-volatile solid-state NAND flash storage connected via PCIe lanes.Stores multi-gigabyte model checkpoint weights (.safetensors) and terabyte-scale training datasets permanently.
Power Supply (PSU) & VRMsConverts AC wall power to low-voltage DC with rapid voltage regulator modules.Modern GPUs exhibit microsecond "transient power spikes" up to 2x rated TDP; underpowered PSUs trip safety shutdowns during heavy matrix workloads.
๐Ÿ’ก Foundational Takeaway: Software cannot run without physical silicon. In AI, performance is rarely dictated by raw Python code syntax; it is dictated by how efficiently data moves between storage, memory, and accelerator compute cores.
02

Compute Engines: CPU, GPU, and NPU Architecture

Moving beyond elementary slogans: understanding latency-optimized vs. throughput-optimized silicon.

A common beginner misconception is: "CPUs are for normal software, GPUs are for 3D video games, and NPUs are for AI." This is inaccurate. Real computing architectures are differentiated by their execution philosophies and memory bandwidth profiles:

Architectural DimensionCPU (e.g. Intel Core / AMD Ryzen)GPU (e.g. NVIDIA RTX 4090 / H100)NPU (e.g. Intel NPU / Apple Neural Engine)
Primary Design GoalMinimize Latency: Finish individual sequential tasks as fast as physically possible.Maximize Throughput: Process millions of independent math calculations concurrently.Maximize Energy Efficiency: High operations per watt for continuous background inference.
Core Architecture8 to 32 large, sophisticated cores with deep branch prediction and out-of-order execution.Thousands of compact SIMD/SIMT cores + dedicated 4th/5th Gen Tensor Cores.Fixed-function systolic arrays and matrix multiplication engines optimized for INT8/FP16.
Clock SpeedsHigh (4.0 GHz to 5.8 GHz).Moderate (1.8 GHz to 2.6 GHz).Moderate to Low (1.0 GHz to 1.8 GHz).
Memory Bandwidth50 โ€“ 90 GB/s (Dual-channel DDR5).1,000 โ€“ 3,350 GB/s (GDDR6X / HBM3e).Shared with system RAM (LPDDR5X, ~80โ€“120 GB/s).
Ideal AI TasksTokenization, regex parsing, pipeline scheduling, token sampling (temperature, top-p).Large-scale model training, backpropagation, batch LLM inference, embedding generation.On-device background AI, eye tracking, real-time noise suppression, local Small Language Models (SLMs).
Power Profile65W โ€“ 250W under full load.150W โ€“ 450W+ (requires liquid or multi-fan cooling).5W โ€“ 15W (ideal for fanless ultrabooks and mobile).

Why AI Workloads Benefit Massively from Accelerator Cores

Consider a standard neural network layer: computing activations requires a massive matrix multiplication:

Mathematical Formulation & Parallel Execution
Y = Activation(W ยท X + B)

Where:
- W (Weight Matrix): 4096 ร— 4096 floats (~16.7 million values)
- X (Input Vector):  4096 floats
- Total Multiply-Accumulate Operations: ~16,777,216 operations for a single layer!

CPU Execution (8 Cores):
- Executes sequentially or in small SIMD vectors (AVX-512).
- Must loop through millions of steps. High clock, but bounded parallelism.

GPU Tensor Core Execution (Thousands of Cores):
- Dispatches matrix tiles across hundreds of Streaming Multiprocessors (SMs).
- Fused Multiply-Add (FMA) instructions compute entire 16ร—16 matrix blocks in a single clock cycle!
โš ๏ธ Software Dependency Note: A GPU cannot automatically accelerate code just because it is installed in your computer. The software must explicitly invoke parallel libraries (such as NVIDIA CUDA, AMD ROCm, or Apple Metal). Ordinary Python loops (e.g. for x in list:) execute strictly on a single CPU core!
03

The Memory Hierarchy, VRAM & The "Memory Wall"

Why memory capacity is a hard wall, memory bandwidth is a speed dial, and the two are never interchangeable.

Computer memory is governed by a fundamental physical constraint: you cannot have memory that is simultaneously ultra-fast, massive in capacity, and inexpensive. Silicon physics dictates that proximity to the compute cores determines speed. Consequently, computers use a hierarchical pyramid:

TierTypeTypical CapacityAccess LatencyTypical BandwidthPersistence
Tier 1CPU Registers< 1 KB per core~0.5 nanoseconds~10,000 GB/sVolatile (Lost on power off)
Tier 2L1 / L2 / L3 SRAM Cache32 KB to 96 MB1 to 15 nanoseconds1,000 to 2,500 GB/sVolatile
Tier 3GPU VRAM (GDDR6X / HBM3e)8 GB to 96 GB10 to 25 nanoseconds1,000 to 3,350 GB/sVolatile
Tier 4Host System RAM (DDR5)16 GB to 128 GB60 to 85 nanoseconds50 to 90 GB/sVolatile
Tier 5NVMe PCIe Gen 4/5 SSD1 TB to 8 TB10 to 50 microseconds7 to 14 GB/sNon-Volatile (Permanent)
Tier 6Cloud Storage / Network (S3)Petabytes5 to 50 milliseconds0.1 to 2.5 GB/sNon-Volatile (Permanent)

The Golden Rule of AI Hardware: "Capacity is a Hard Wall, Bandwidth is a Speed Dial"

In AI Engineering, understanding the difference between memory capacity and memory bandwidth is essential:

Capacity vs. Bandwidth Dynamics
๐Ÿงฑ 1. Memory Capacity (Gigabytes) โ€” The Hard Wall
Determines whether a model can run at all. You must store: (Static Model Weights) + (Dynamic KV Cache) + (Runtime Buffers). If total footprint exceeds VRAM by even 100 MB, the process crashes with CUDA out of memory or drops into CPU offloading.
โšก 2. Memory Bandwidth (GB/s) โ€” The Speed Dial
Determines how fast tokens are generated once the model fits. Because autoregressive decoding streams all parameters through the GPU for every single token, an 8B model on a 300 GB/s card generates ~15 tokens/s, while on a 1,000 GB/s card it generates ~45 tokens/s.

Dedicated VRAM vs. Unified Memory Architectures

Why can't you just add 16 GB System RAM and 8 GB GPU VRAM together to get 24 GB?

In standard PC architecture (Windows/Linux with discrete GPUs), the GPU and CPU have physically isolated memory pools connected only by the PCIe expansion bus. The GPU cannot access DDR5 System RAM at native VRAM speedsโ€”it must request transfers across the PCIe bus, which is 30x to 60x slower.

In contrast, Unified Memory Architectures (such as Apple Silicon M-series chips or high-end server APUs like NVIDIA GH200) place the CPU and GPU on the same substrate, sharing a single high-bandwidth LPDDR5X memory pool (up to 128GBโ€“192GB at 400โ€“800 GB/s). While peak bandwidth is lower than dedicated HBM3e, it eliminates the PCIe bus barrier and allows running massive 70B+ models without multi-GPU clusters.

04

Data Representation & Numerical Precision in AI

From raw binary voltage switches to floating-point tensors and INT4 quantization.

Computers operate exclusively on binary digits (bits: 0 or 1), physically represented by microscopic transistor voltage levels. Eight bits constitute a Byte (256 distinct values). From bytes, we construct hexadecimal notations, ASCII/Unicode text encodings, and numerical data types.

In AI Engineering, every neural networkโ€”whether processing natural language, images, or audioโ€”converts inputs into tensors (multidimensional arrays of numerical values). The amount of memory a model requires is directly determined by its numerical precision:

Precision FormatBits per ParameterBytes per Parameter8B Model Weight Size70B Model Weight SizeIndustry Use Case
FP32 (Single Precision)32 bits4 bytes~32.0 GB~280 GBHistorical baseline; legacy scientific computing.
FP16 / BF16 (Half Precision)16 bits2 bytes~16.0 GB~140 GBModern Standard: Model training and high-fidelity server inference.
INT8 (8-Bit Integer)8 bits1 byte~8.0 GB~70 GBStandard quantization for enterprise production serving with minimal perplexity loss.
INT4 / FP4 (4-Bit Quantized)4 bits0.5 bytes~4.5 GB~38 GBConsumer Deployment: Enables running 8B models on 8GB laptops and 70B models on workstations.
๐Ÿ’ก Fast Memory Sizing Formula: To calculate the minimum VRAM required just to store an AI model's weights:
VRAM (GB) โ‰ˆ Parameters (Billions) ร— Bytes per Parameter ร— 1.2 (for runtime & KV cache overhead)
Example for Llama-3 8B at FP16: 8 ร— 2 ร— 1.2 โ‰ˆ 19.2 GB. At INT4: 8 ร— 0.5 ร— 1.2 โ‰ˆ 4.8 GB.
05

Motherboard, PCIe Lanes & Data Movement Bottlenecks

How components communicate and why the PCIe bus is the most critical bottleneck in hardware scaling.

The Motherboard is the physical highway system connecting CPU, memory, accelerators, and storage via etched copper traces and bus protocols:

Bus / InterconnectGenerational StandardMax Bandwidth (x16 Lanes)Impact on AI Engineering
PCIe 4.0Mainstream Consumer~31.5 GB/sHost-to-Device transfer bottleneck; takes ~0.5s to transfer 16GB model weights.
PCIe 5.0Modern High-End / Server~63.0 GB/sDoubles transfer rate; reduces GPU weight loading and layer offloading latency.
PCIe 6.02026 Enterprise / AI Clusters~128.0 GB/s (PAM4 / FLIT)Essential for multi-GPU interconnects and 800G optical networking clusters.
NVLink (NVIDIA Proprietary)Data Center (H100 / Blackwell)900 โ€“ 1,800 GB/sAllows multiple GPUs to pool memory seamlessly as if on a single giant chip.
๐Ÿšจ The PCIe Offloading Performance Cliff: When your model exceeds VRAM, frameworks attempt to "offload" layers into System RAM. While the GPU can read its own VRAM at 1,000 GB/s, accessing the offloaded layers across a PCIe 4.0 slot occurs at only 31.5 GB/s (a 97% bandwidth drop!). This is why offloaded models crawl from 35 tokens/s down to 1โ€“2 tokens/s.
06

What Actually Happens When You Run Python / AI Code?

The chronological journey from typing a command in your terminal to outputting tokens.

When you open your terminal and execute python run_inference.py, an orchestrated sequence of events occurs across the operating system and physical hardware:

Chronological Execution Lifecycle
1. OS Process Creation
OS allocates a unique Process ID (PID), virtual address space, and loads the Python interpreter binary into Host RAM.
2. Bytecode Compilation
Python source code is parsed into an Abstract Syntax Tree (AST) and compiled into Python bytecode executed by the CPython runtime.
3. Model Weight Streaming (Disk โ†’ RAM)
PyTorch requests multi-gigabyte weight files (e.g. model.safetensors) from the NVMe SSD into Host System RAM buffers.
4. PCIe Host-to-Device Transfer
CUDA runtime uses Direct Memory Access (DMA) to stream weights from System RAM across the PCIe bus into GPU VRAM.
5. Tensor Kernel Execution on GPU
CUDA launches thousands of parallel GPU threads. Tensor Cores execute matrix multiplications, writing activations back to VRAM.
6. Result Extraction & Display
Resulting logits are transferred back across PCIe to System RAM, decoded into Unicode text tokens, and displayed to stdout.
07

Interactive Lab: AI Hardware Workload Explorer

Simulate how different software workloads stress CPU, GPU, VRAM, and storage subsystems.

Genuinely Interactive Educational Simulator* Illustrative simulation for hardware reasoning

AI Hardware Stress & Bottleneck Simulator

Select a computational workload below, tweak the simulated hardware configuration, and observe real-time subsystem utilization and bottleneck diagnoses.

1. Select Active Workload
2. Customize Simulated Machine Hardware
8 Cores
16 GB
8 GB
Mid-Tier (RTX 4070)
NVME-GEN4
1x Scale
3. Real-Time Subsystem Stress Gauges
CPU Load64%
8 Cores
System RAM16 / 16 GB
90% Used
GPU VRAM18 / 8 GB
100% Used
Storage Read I/O7% Saturation
500 MB/s
Estimated ThroughputSimulation
0.8 - 1.5 tokens/s (PCIe offloaded)
CRITICAL: CUDA Out of Memory (VRAM Exceeded)
The workload requires ~18 GB of VRAM, but this machine only has 8 GB. Layers will spill to System RAM across the PCIe bus or crash immediately with an OOM error.
08

Interactive Tool: Architecture & Data-Flow Simulator

Step through the journey of tensors across SSD, RAM, PCIe, VRAM, and GPU Tensor Cores.

Hardware Pipeline Stage: 1. SSD Storage Read

Operating system loads model weights (.safetensors) from NVMe SSD.

๐Ÿ’พ
SSD Storage Read
5,000 - 14,000 MB/s
โ†’
๐Ÿ”‹
Host RAM Staging
50 - 90 GB/s
โ†’
๐Ÿ›ฃ๏ธ
PCIe Bus Transfer
31.5 - 63.0 GB/s
โ†’
๐Ÿ’Ž
VRAM Allocation
1,000 - 3,350 GB/s
โ†’
๐ŸŽฎ
Tensor Core Compute
1.8 - 2.6 GHz
โ†’
๐Ÿง 
Output Sampling
4.0 - 5.7 GHz
๐ŸŽฎComponent Deep Inspector: Graphics Processing Unit (GPU / Tensor Cores)
Hardware Speed / Bandwidth
1.8 - 2.6 GHz
Access Latency
Throughput oriented
Typical Capacity
4,000 - 18,000 Cores
Physical Role
Massively parallel compute engine executing thousands of arithmetic matrix operations simultaneously.
Hardware Limitation
High power draw (150-450W). Inefficient at branchy sequential code or OS task scheduling.
AI Engineering Significance
The workhorse of deep learning. Tensor Cores execute FP16/BF16/INT4 matrix multiply-accumulate operations.
09

Real-World Hardware Debugging Challenge

Troubleshoot an actual production incident: an AI application running slowly and crashing under load.

Production Incident Scenario #814

Incident: Local 8B LLM Service Crashing & Crawling at 1.2 Tokens/sec

Scenario Description: An engineering team deployed a customer support bot using an 8-billion parameter open-weights model on an on-premise development workstation. Originally, when tested with a tiny prompt (512 tokens), it generated responses at 32 tokens/second. However, after expanding the document context window to 8,000 tokens for RAG retrieval, the application crashed with:

torch.cuda.OutOfMemoryError: CUDA out of memory. 
Tried to allocate 16.42 GiB (GPU 0; 8.00 GiB total capacity; 6.12 GiB already allocated; 
1.88 GiB free; 6.15 GiB reserved in total by PyTorch)

When the developer enabled automatic CPU layer offloading (device_map="auto"), the crash stopped, but generation speed collapsed to an unusable 1.2 tokens/second.

Fictional Machine Specifications
Processor (CPU)
8-Core AMD Ryzen 7 7700X (4.5 GHz)
System RAM
32 GB DDR5-5600 MT/s
Graphics (GPU)
NVIDIA RTX 4060 (8 GB GDDR6)
PCIe Interconnect
PCIe 4.0 x8 Slot (~15.7 GB/s)
Storage Drive
1 TB NVMe PCIe Gen 4 SSD
Model Precision
FP16 (16-Bit Float, 2 Bytes/param)
Step 1: Identify the Root Cause Hardware Bottleneck
10

Notes, Bottleneck Matrix & What to Learn Next

The AI engineer's quick-reference matrix and transition to the next curriculum milestone.

Hardware Bottleneck Quick Diagnostic Matrix

Symptom ObservedUnderlying Hardware BottleneckCorrect Remediation Strategy
CUDA out of memory error on boot or batch launchVRAM Capacity exceeded (Weights + KV Cache > GPU VRAM)Apply 4-bit/8-bit quantization; reduce batch size; shorten context length; deploy tensor parallelism across multiple GPUs.
Tokens generate at < 2 tokens/s with low GPU compute utilizationPCIe Bus Bandwidth starvation due to layer offloadingFit the model entirely inside VRAM using quantization or upgrade to a higher-capacity GPU; avoid PCIe offloading.
GPU compute utilization drops to 0% periodically during trainingStorage Read I/O or CPU DataLoader starvationMove datasets to an NVMe PCIe Gen 4 SSD; increase PyTorch num_workers; cache preprocessed tensors in RAM.
System freezes and disk activity LED stays pinned at 100%System RAM Out-of-Memory (OS Swap Thrashing)Increase system RAM; use memory-mapped tensors (mmap); stream large datasets from disk using generators.

What to Learn Next

Now that you understand how physical silicon, memory hierarchies, and accelerators function, the next logical question is: who manages all this hardware and shares it fairly among running applications?

That is the job of the Operating System. In the next topic, you will explore how the OS kernel manages virtual memory, schedules threads across CPU cores, communicates with GPUs via device drivers, and coordinates processes.

What You Should Know Now (Competency Checklist)

Verify your mastery of Computer Fundamentals before proceeding to Operating Systems. Check each competency as you confirm your understanding:

I understand the layered mental model: Application โ†’ OS/Drivers โ†’ Accelerators (CPU/GPU/NPU) โ†’ Memory Subsystem โ†’ Storage.
I can explain why LLM autoregressive token generation is memory-bandwidth bound rather than compute bound.
I know the difference between CPU out-of-order latency optimization vs GPU massive parallel SIMD throughput.
I know what an NPU does and why 40-60+ TOPS at 5-15W matters for on-device AI laptops.
I understand the memory hierarchy and why DDR5 System RAM and GPU VRAM cannot simply be summed together.
I can calculate the memory footprint of an AI model using parameter count and precision (FP32, FP16, INT8, INT4).
I know why moving data across the PCIe bus creates a severe bottleneck for offloaded AI workloads.
I can systematically identify whether an AI application bottleneck is caused by VRAM capacity, memory bandwidth, or I/O starvation.
๐Ÿงช Knowledge Assessment QuizQuestion 1 of 8
In the context of Large Language Model (LLM) inference, why is autoregressive token generation (the decode phase) strictly memory-bandwidth bound rather than compute-bound?