Docker Deep Dive · Prerequisites · P5 of 7

Kernel, Syscalls & Namespaces

One program owns your machine: the kernel. Everything else — your shell, your apps, every container — has to ask it for favors. Today you watch the asking happen, then use it to build a piece of a container with your bare hands.

⏱ ~30 min hands-on · repo chapter p5-kernel-syscalls-namespaces · cheatsheet: Linux & Shell

prereq P5 / 7

Before you start

cd docker-deep-dive
git pull
cd p5-kernel-syscalls-namespaces
docker build -t oslab .

That build bakes strace — today's x-ray machine — into an Ubuntu image called oslab. If you've done Lesson 3 you know exactly what just happened; if you're doing prerequisites first, fine: it downloads a tool into a snapshot so the labs below can use it. Lesson 3 explains the machinery.

1.The wall

Your machine's memory is split into two worlds. In user space live ordinary programs: your shell, ls, Python, Postgres — powerless on their own. In kernel space lives the one program with real authority: the kernel, which owns the processes (P4), the memory, the filesystems (P2), and the network (P6). A user-space program cannot open a file, start a process, or send a byte by itself. It has to ask.

The ask is a system call — syscall — the only doorway through the wall. There are a few hundred of them (openat, read, write, execve, socket, clone…), and every single thing every program does eventually funnels through one.¹ Programs rarely dial them directly — the C library (libc) wraps them in friendly functions — but the gate is the gate.

USER SPACE — ordinary programs, no authority your shell ls · cat python every container libc — friendly wrappers around the raw syscalls openat read write execve socket the wall → KERNEL SPACE — the one program with authority processes P4's PIDs, signals memory pages, OOM killer filesystems P2's tree, permissions network P6's sockets, ports
Two worlds, one wall, a few hundred gates. User-space programs — containers included — have no direct power; every action becomes a syscall through the wall. Keep this picture: the rest of the lesson is just walking through those gates and reading the kernel's guest book.

2.Lab A — one kernel, many userlands

All your containers share one Linux kernel — on macOS and Windows it lives in Docker Desktop's hidden VM; on a Linux host it's the host's own. You can now prove it from the inside. uname -r asks the kernel to identify itself:

docker run --rm ubuntu:24.04 uname -r
docker run --rm alpine uname -r
5.10.76-linuxkit
5.10.76-linuxkit

Ubuntu container and Alpine container: identical kernel, byte for byte — on the instructor's machine, the linuxkit kernel of Docker Desktop's VM. Now ask your host the same question and watch the pretence show (each row is that OS's own spelling — output shown for orientation, answers will differ per machine):

your hostask ittypical answer
macOS Terminaluname -srDarwin 25.5.0 — not Linux at all; the shared kernel lives in the VM
Windows PowerShellcmd /c verMicrosoft Windows … — same story, the kernel lives in WSL 2
Linuxuname -srLinux 6.8.0-… — the very kernel your containers just reported

So what's actually "Ubuntu" about an Ubuntu container? Only the files — P2's tree: its own /etc, its own apt, its own libc version:

docker run --rm ubuntu:24.04 head -2 /etc/os-release
docker run --rm alpine head -2 /etc/os-release
PRETTY_NAME="Ubuntu 24.04.4 LTS"
NAME="Ubuntu"
NAME="Alpine Linux"
ID=alpine

A Linux distribution is a kernel plus a userland — and in containers, the kernel is factored out and shared.² That's why an "operating system" image can be 8 MB (Alpine): it ships no kernel at all, just user-space files.

Why this matters for Docker & Kubernetes

Small images, instant starts, and dense packing all fall out of kernel-sharing — but so does the blast radius: a kernel panic, or a kernel that's too old for your app's syscalls, affects every container on the node. When Lesson 9's multi-arch section said images are per-CPU-architecture, this is why: user-space binaries must speak the shared kernel's architecture.

3.Lab B — strace, the x-ray machine

Bare Ubuntu can't show you syscalls — try it and meet a P4 exit-127-style error at the docker level:

docker run --rm ubuntu:24.04 strace ls
docker: Error response from daemon: … exec: "strace": executable file not found in $PATH: unknown.

Hence the oslab image you built. Enter it and x-ray a command you've used since P1:³

docker run -it --rm oslab bash
strace -c ls /   # -c: run it, then tally every syscall it made
% time     seconds  usecs/call     calls    errors syscall
------ ----------- ----------- --------- --------- ----------------
 24.00    0.000360          25        14           mmap
 20.80    0.000312          39         8           close
  8.40    0.000126          18         7           fstat
  5.73    0.000086          14         6           openat
  4.53    0.000068          34         2           getdents64
  4.07    0.000061          12         5           read
  1.20    0.000018          18         1           write
  0.00    0.000000          0          1           execve

Little ls crossed the wall ~60 times: execve to become ls at all (P4's exec!), openat to reach the directory, getdents64 to read its entries, one write to print — P3's stdout, gate-level. Now narrow the x-ray to just the file-opens:

strace -e trace=openat cat /etc/hostname
openat(AT_FDCWD, "/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/lib/aarch64-linux-gnu/libc.so.6", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/etc/hostname", O_RDONLY) = 3
01660e73a23e
+++ exited with 0 +++

Read it like a story: cat loads its libc, opens your file, and the kernel answers each request with a small number — = 3 — a file descriptor, the very same numbers P3 taught you as 0, 1, 2. Then the hostname prints (it's the container ID — Lesson 1's UTS fact), and the process exits 0 — P4's exit code, reported by strace itself.

Who's allowed to x-ray? (a Lesson 11 preview)

strace works via the ptrace syscall — spying on another process is security-sensitive. Docker filters every container's syscalls through a seccomp profile (a syscall firewall in the kernel); ptrace has been on the default allow-list since Docker 19.03, which is why this worked with no extra flags.⁴ On older engines or hardened setups you'd add --cap-add SYS_PTRACE. Lesson 11 turns these dials the other way — removing capabilities.

4.Watch one command cross the wall

Here is that cat /etc/hostname trace again — as a film. Every prerequisite lesson you've done is one of the frames:

USER SPACE cat wants /etc/hostname your terminal stdout, fd 1 (P3) the gate KERNEL SPACE libc: openat("/etc/hostname") CPU switches to kernel mode path walk: /etc/hostname owner? group? other? → r-- P2's permission check ✓ = 3 (a file descriptor) read(3, …) → bytes "01660e73a23e" write(1, "01660e73a23e") — P3's stdout wire exit 0 P4's $? = 0
  1. cat asks libc to open the file; libc prepares the raw openat syscall — the request approaches the gate.
  2. The CPU itself switches from user mode to kernel mode. This is the only legal way across the wall.
  3. Inside, the kernel walks the path and runs P2's exact permission algorithm — owner? group? other? — before agreeing.
  4. The kernel hands back a small number, fd 3 — P3's file descriptors were these all along.
  5. read(3, …) crosses again: the kernel copies the file's bytes into cat's memory.
  6. write(1, …) sends them out fd 1 — stdout — which your terminal is holding, exactly as P3 drew it.
  7. exit_group(0): the process ends and the shell collects P4's exit code. Every command you've ever run did all of this.
P2 + P3 + P4, reunited at the gate. Permissions are checked in the kernel, streams are kernel-managed descriptors, exit codes are collected by the kernel — the prerequisite phase has been describing kernel behavior all along.

5.Lab C — /proc: the kernel, live, as files

P2 promised "everything is a file" gets weird and wonderful. Here it is: /proc is a fake filesystem the kernel synthesizes on the fly — its live internal state, dressed up as readable files.⁵ One numbered directory per process:

ls /proc | head -14
tr "\0" " " < /proc/1/cmdline; echo   # what is PID 1 here, really?
ls -l /proc/self                      # /proc/self = "whoever is asking"
head -3 /proc/meminfo
1
7
8
buddyinfo
bus
cgroups
cmdline
consoles
cpuinfo
…
bash                                   ← PID 1 is the bash you asked docker to run
lrwxrwxrwx 1 root root 0 Jul 17 03:16 /proc/self -> 9   ← points at the ls that asked!
MemTotal:        6081840 kB
MemFree:          249840 kB
MemAvailable:    2184212 kB

Three things worth savoring. The numbered dirs are the process list — and in this container there are only a couple, P4's tiny tree. /proc/self is a symlink that points at whichever process reads it. And MemTotal says ~6 GB — on Docker Desktop that's not your machine's RAM, it's the VM's allowance (a Linux host would show its real total); the kernel answering is the one your containers share.

Now the punchline: ps is not magic — it just reads /proc. X-ray it:

strace -e trace=openat ps aux 2>&1 | grep "/proc" | head -8
openat(AT_FDCWD, "/proc/self/stat", O_RDONLY) = 3
openat(AT_FDCWD, "/proc/uptime", O_RDONLY) = 3
openat(AT_FDCWD, "/proc/12/status", O_RDONLY) = 3
openat(AT_FDCWD, "/proc/12/stat", O_RDONLY) = 3
openat(AT_FDCWD, "/proc", O_RDONLY|O_NONBLOCK|O_CLOEXEC|O_DIRECTORY) = 3
openat(AT_FDCWD, "/proc/uptime", O_RDONLY) = 4
openat(AT_FDCWD, "/proc/1/stat", O_RDONLY) = 4
openat(AT_FDCWD, "/proc/1/status", O_RDONLY) = 4

Every P4 tool decomposes this way: ps reads /proc, docker stats reads cgroup files, ss reads /proc/net. Tools are just polite formatting over kernel files.

/proc/PID/environ — why env vars aren't secrets

A process's environment (P4) is readable at /proc/PID/environ:

docker run --rm -e DB_PASSWORD=hunter2 ubuntu:24.04 sh -c 'tr "\0" "\n" < /proc/1/environ'
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
HOSTNAME=94dd9c883ef9
DB_PASSWORD=hunter2
HOME=/root

Anything that can read that file has your password — the kernel-level reason Lesson 4 said docker inspect shows env in plain text, and Lesson 11 moved secrets to files.

6.Lab D — namespaces, by hand

Lesson 10 showed namespaces through Docker's flags. Here's the raw material. Which namespaces does your process live in? The kernel lists your memberships — as files, of course:

ls -l /proc/$$/ns   # $$ = this shell's PID (P4)
lrwxrwxrwx 1 root root 0 Jul 17 03:16 cgroup -> cgroup:[4026532894]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 ipc -> ipc:[4026532815]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 mnt -> mnt:[4026532813]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 net -> net:[4026532818]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 pid -> pid:[4026532816]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 user -> user:[4026531837]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 uts -> uts:[4026532814]

Each line is a room this process is locked into; the number identifies the room. Processes with the same number share that view of the world.⁶

namespaceisolatesyou met it in
pidwhich processes existLesson 10's PID 1994-vs-PID 1
netinterfaces, IPs, portswhy -p exists (Lesson 4, P6 next)
mntthe filesystem viewimage layers become / (Lessons 2–3)
utshostnamecontainer ID in the prompt (Lesson 1)
ipcshared-memory channels(background — rarely touched by hand)
useruid/gid mappingnote its low number above: shared with the host — Docker doesn't unshare it by default, Lesson 11's rootless story

Now build one. unshare is a small wrapper around the syscall of the same name: "put my next command in new rooms".

docker run -it --rm --privileged ubuntu:24.04 bash
unshare --pid --fork --mount-proc bash   # new PID namespace + fresh /proc
ps aux
exit   # leave the namespace…
exit   # …then the container
USER       PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root         1  0.0  0.0   4032  1272 ?        S    03:16   0:00 bash
root         2  0.0  0.0   7628  3584 ?        R    03:16   0:00 ps aux

Look at that ps: your bash is PID 1, and the world contains two processes. That's Lesson 1's "sandboxed process" punchline and Lesson 10's Lab A — reproduced with one command and zero Docker. You just built one-seventh of a container. Docker's actual job is assembling all the rooms at once — plus cgroups, plus the image's mnt-namespace filesystem, plus a nice API. "An isolated, restricted process", nothing more mystical than that.²

--privileged is a learning-lab flag

Creating namespaces inside a container needs powers Docker rightly strips by default (CAP_SYS_ADMIN among them⁷). --privileged hands the container essentially the Docker host's full power — fine for a throwaway experiment on your own machine, a serious finding anywhere near production. Lesson 11 is the antidote chapter.

7.cgroups — the other half, briefly

Namespaces limit what a process can see. The kernel's other lever, cgroups, limits what it can use — memory, CPU, I/O. Same design language: the interface is files.⁸

docker run --rm --memory=64m ubuntu:24.04 cat /sys/fs/cgroup/memory.max
67108864

64 MiB in bytes. Docker's --memory flag did nothing clever — it wrote that number into a cgroup-v2 file, and the kernel enforces it. When a process pushes past the ceiling, the kernel's OOM killer ends it with SIGKILL — which P4 taught you to read in the exit code: 128 + 9 = 137. Lesson 10 stages that whole crime scene deliberately; now you know both halves of its machinery.

You can now

8.Check yourself

A system call is best described as…

  1. a library call inside libc
  2. the doorway into the kernel
  3. a process the kernel starts
  4. an interrupt the shell handles

Namespaces vs cgroups — which split is right?

  1. namespaces limit seeing; cgroups limit using
  2. namespaces limit using; cgroups limit seeing
  3. namespaces isolate files; cgroups isolate networks
  4. namespaces restrict memory; cgroups restrict processes

An Ubuntu container and an Alpine container on your machine print the same uname -r. Why?

  1. Docker copies the kernel into images
  2. all containers share one host kernel
  3. uname is cached by the daemon
  4. both images ship identical kernel builds

What is /proc, and name one thing you'd read from it while debugging.

A fake filesystem the kernel synthesizes: its live state exposed as files, one numbered directory per process. Debug reads: /proc/1/cmdline (what PID 1 really is), /proc/PID/environ (a process's env — why secrets leak), /proc/meminfo (the VM's real memory), /proc/PID/ns (namespace memberships).

You ran unshare --pid --fork --mount-proc bash and then ps aux. What did ps show, and why?

Two processes, with bash as PID 1: unshare created a new PID namespace (bash renumbered to 1, everything else invisible) and mounted a fresh /proc so ps — which just reads /proc — sees only the new namespace's processes. That's the core of what Docker does when it starts a container.

9.Go deeper

Primary source: Ivan Velichko's Learning Containers From The Bottom Up — the "isolated + restricted process" mental model this lesson builds toward, with hands-on follow-ups on iximiuz labs. For strace joy, Julia Evans' strace posts & zine; for the dry truth, man7's namespaces(7) and proc(5).

Next: the net namespace got one row in a table — it deserves its own lesson. P6: Ports, DNS & Packages opens the network side: sockets, the 127.0.0.1-vs-0.0.0.0 trap behind every "why can't I reach my container", and where apt actually puts things. And Lesson 10, when you reach it later in the course, will read like a friendly rerun.

Stuck? Curious?

Bring questions to class, or open an issue on the course repo — include the command you ran and the output you got. The quizzes above are for self-checking: commit to an answer before revealing it, and re-try anything you missed tomorrow.

Sources

  1. man7 — syscalls(2) (what syscalls are; the full list)
  2. Ivan Velichko — Learning Containers From The Bottom Up (kernel+userland; isolated & restricted process)
  3. Julia Evans — strace posts & zine (strace technique)
  4. Docker Docs — seccomp profiles (syscall filtering; ptrace allowed since 19.03)
  5. man7 — proc(5) (/proc structure, environ, ns)
  6. man7 — namespaces(7) (namespace kinds and /proc/PID/ns)
  7. man7 — capabilities(7) (CAP_SYS_ADMIN and friends)
  8. man7 — cgroups(7) (cgroup v2 file interface)