status: all systems learnable

You don't need to
be "a math person" to
run infrastructure.

DevOps math is narrow and repetitive: a handful of ideas — binary, ratios, percentages, percentiles — show up again and again in subnetting, SLAs, dashboards, and capacity planning. Learn these seven, and nothing in the course will surprise you.

binary / hex CIDR 99.9% uptime GB vs GiB p95 / p99 req/sec cron
01

Number systems: binary, octal, hex

Why it's here: IP addresses, subnet masks, file permissions (chmod 755), color codes, and memory addresses are all just decimal numbers written in a different base.

A "base" is just how many digits you get before you carry over. Decimal (base 10) is what you grew up with. Computers prefer binary (base 2) because a transistor is either on or off. Humans reading binary get tired of counting 1s and 0s, so we compress it into hex (base 16) — every 4 binary digits become exactly 1 hex digit.

  • Binary — digits 0–1. Each position is a power of 2 (…32, 16, 8, 4, 2, 1).
  • Octal — digits 0–7. Used in Linux file permissions, e.g. chmod 755.
  • Hex — digits 0–9 then A–F. Used in MAC addresses, IPv6, color codes, memory dumps.
Trick for binary→decimal: write the powers of 2 above each bit, then add up the ones where the bit is 1. 1011 → 8+0+2+1 = 11.
base-converter.sh
Quick check — what does binary 1010 equal in decimal?
02

Subnetting & CIDR notation

Why it's here: every VPC, security group, and Kubernetes cluster network you touch is defined by a CIDR block like 10.0.0.0/24. Knowing what the /24 means is non-negotiable.

An IP address like 192.168.1.10 is really 32 bits (4 groups of 8, called octets). The /24 in 192.168.1.0/24 says: "the first 24 bits are fixed (the network), the remaining 8 bits are free (the hosts)." Fewer bits reserved for hosts means fewer usable addresses, but more subnets.

  • Prefix (/n) — number of bits locked as "network."
  • Usable hosts — 2(32-n) − 2 (one address is reserved for the network, one for broadcast).
  • /32 = a single host. /0 = the entire internet.
subnet-calc.sh
A /28 network has how many usable host addresses?
03

Uptime & SLA percentages

Why it's here: "five nines" isn't a slogan, it's a strict downtime budget. SREs and DevOps engineers price incidents against exactly this math.

An SLA of 99.9% uptime means the system is allowed to be down for 0.1% of the time. The trap: 0.1% of a year sounds tiny, but a year is 525,600 minutes — so 0.1% is still ~8.8 hours. Each extra "nine" cuts the allowed downtime by roughly 10x.

sla-budget.sh — the signature calc
uptimedowntime allowed
Roughly how much downtime per year does 99.99% ("four nines") allow?
04

Storage units: bytes to petabytes

Why it's here: provisioning a volume, reading a Prometheus disk-usage graph, or sizing an S3 bucket all depend on converting between units — and on the classic GB vs GiB trap.

Storage scales in powers of 1024 (210), not 1000 — because computers count in binary. So technically 1 KiB = 1024 bytes, while 1 KB is sometimes (loosely) used for 1000 bytes. Most tools you'll use (df, du, cloud consoles) mix these conventions, which is exactly why unit math trips people up in practice.

unit-convert.sh
How many MB are in 1 GB (binary convention, ×1024)?
05

Statistics for monitoring: mean, median, percentiles

Why it's here: every latency dashboard talks about "p95" or "p99" response time. That's a percentile — and it tells a very different story than an average.

The average (mean) can hide a bad outlier. If 99 requests take 10ms and 1 takes 5000ms, the average looks fine, but that one slow user had a terrible time. A percentile answers: "what value is X% of requests faster than?" p95 = 200ms means 95% of requests finished in 200ms or less — and the worst 5% took longer. DevOps teams watch p95/p99, not the mean, because those tails are where users actually feel pain.

latency-stats.sh
A dashboard shows p99 latency = 900ms while the mean is 40ms. What does that tell you?
06

Rate & throughput math

Why it's here: capacity planning, autoscaling thresholds, and load-testing reports all come down to "how many things per second" and "how long is the queue."

Throughput is just division: requests ÷ seconds = requests per second (RPS). To plan capacity, flip it around: if one server handles 200 RPS and you expect 5,000 RPS at peak, you need at least 5000 ÷ 200 = 25 servers (plus headroom for failures).

A related idea, Little's Law, connects queueing: average items in the system = arrival rate × average time in system. If requests arrive at 50/sec and each takes 0.4s to process, on average 20 requests are "in flight" at once — that's how many concurrent connections/threads you need to provision for.

capacity-calc.sh
Requests arrive at 100/sec, each takes 0.2s to process. By Little's Law, how many are "in flight" on average?
07

Cron expressions & time math

Why it's here: scheduled pipelines, backup jobs, and cert-renewal scripts are all triggered by cron syntax — five fields of modular arithmetic in disguise.

A cron expression has 5 fields, each with its own valid range — this is just modular arithmetic (numbers that wrap around, like a clock):

reference.txt
fieldrangeexample
minute0–59*/15 → every 15 min
hour0–232 → 2am
day of month1–311 → the 1st
month1–12*/3 → quarterly
day of week0–6 (Sun=0)1-5 → weekdays
0 2 * * * reads as: minute 0, hour 2, every day, every month, every weekday → "run at 2:00am, daily."
What does */15 * * * * mean?