epoch.training
A supercomputer at a national laboratory

$ ./epoch-training --whoami

The machines that train AI
/ explained by someone who ran them.

Supercomputing and AI-infrastructure training for engineers who want to stop treating the cluster like a black box. Taught by someone who architected top-100 supercomputers, handled storage and configuration management behind hurricane forecasts, and helped run the grid that confirmed the Higgs boson.

// the whole discipline is one question: correct, or fast?  the answer is yes — and the engineering that makes "both" possible is the entire subject.

scroll ↓

Track record

kris@epoch — machines I've actually run
$ whoami --career
30 years Unix · 19 years HPC · currently: datacenter-GPU performance @ AMD
BioHive-1TOP500 #84 — architected & built in 90 days, $1.6M under budget
LHC Computing Gridhelped run the grid that confirmed the Higgs boson (CMS/ATLAS)
NERSCHopper (150K+ cores), Edison — DOE computational science
CSCSPiz Daint — #3 on the TOP500 for years
NOAA / FNMOCstorage & configuration management behind hurricane-forecast & naval weather models
storage SMELustre & VAST — petabyte-scale parallel filesystems, at real load
$ _

Selected publications

Peer-reviewed, at the venues that matter.

Not a marketing claim — published work on scheduling and petascale storage at the Cray User Group and the Parallel Data Storage Workshop (SC).

  1. S. Alam, H. El-Harake, K. Howard, N. Stringfellow, F. Verzelloni
    6th Parallel Data Storage Workshop (PDSW '11), held with SC11 · 2011 · CSCS
  2. G. Renker, N. Stringfellow, K. Howard, S. Alam, S. Trofinoff
    Cray User Group (CUG) · 2011 · CSCS

What you'll learn

How the machines actually work — not how the vendor slide says they do.

The training splits along the axis every HPC engineer eventually meets: the half that keeps the job correct, and the half that makes it fast. Real systems need both.

// stability

Storage & the I/O wall

Parallel filesystems (Lustre, VAST), why 10,000 ranks hitting one disk is its own discipline, and why the storage decision quietly decides the whole run.

// speed

GPUs & the roofline

Why AI lives on GPUs, memory-bound vs compute-bound, and why your accelerator is "fast" and still 90% idle.

// stability

The network & MPI

Interconnects, collectives, and why a $200M machine spends most of its time waiting on the fabric — not computing.

// speed

Scaling & its limits

Amdahl's law, "just add nodes" and where it breaks, and how to find the bottleneck you actually have instead of the one you assumed.

// stability

Reliability at scale

Checkpointing, node failure, silent data corruption — everything that goes wrong at hour 47, and how the pros plan for it.

// speed

Precision & performance

FP8/FP16/FP32, mixed precision, and the throughput-vs-correctness tradeoff that can quietly ruin a model.

Trainings

Workshops & cohorts

Dates for the next cohorts are being scheduled. Join a waitlist and you'll be first to hear — before it opens publicly.

Loading trainings…

Free · no course to buy first

The Supercomputing Field Guide

Twenty concepts every infrastructure engineer should understand about the machines AI runs on — one page each, in plain language. It's the on-ramp to the full training. Drop your email and I'll send it.

One email with the guide, then occasional field notes on HPC/AI infrastructure. No spam. Unsubscribe anytime.