Docker Deep Dive · Prerequisites · P5 of 7
Kernel, Syscalls & Namespaces
One program owns your machine: the kernel. Everything else — your shell, your apps, every container — has to ask it for favors. Today you watch the asking happen, then use it to build a piece of a container with your bare hands.
Before you start
cd docker-deep-dive
git pull
cd p5-kernel-syscalls-namespaces
docker build -t oslab .
That build bakes strace — today's x-ray machine — into an Ubuntu image
called oslab. If you've done Lesson 3 you know exactly what just happened;
if you're doing prerequisites first, fine: it downloads a tool into a snapshot so the labs
below can use it. Lesson 3 explains the machinery.
1.The wall
Your machine's memory is split into two worlds. In user space live
ordinary programs: your shell, ls, Python, Postgres — powerless on their own.
In kernel space lives the one program with real authority: the
kernel, which owns the processes (P4), the memory, the filesystems (P2),
and the network (P6). A user-space program cannot open a file, start a process, or send a
byte by itself. It has to ask.
The ask is a system call — syscall — the only doorway through
the wall. There are a few hundred of them (openat, read,
write, execve, socket, clone…), and
every single thing every program does eventually funnels through one.¹
Programs rarely dial them directly — the C library (libc) wraps them in
friendly functions — but the gate is the gate.
2.Lab A — one kernel, many userlands
All your containers share one Linux kernel — on macOS and Windows it lives in
Docker Desktop's hidden VM; on a Linux host it's the host's own. You can now prove it
from the inside. uname -r asks the kernel to identify itself:
docker run --rm ubuntu:24.04 uname -r
docker run --rm alpine uname -r
5.10.76-linuxkit
5.10.76-linuxkit
Ubuntu container and Alpine container: identical kernel, byte for
byte — on the instructor's machine, the linuxkit kernel of Docker Desktop's
VM. Now ask your host the same question and watch the pretence show (each row is
that OS's own spelling — output shown for orientation, answers will differ per machine):
| your host | ask it | typical answer |
|---|---|---|
| macOS Terminal | uname -sr | Darwin 25.5.0 — not Linux at all; the shared kernel lives in the VM |
| Windows PowerShell | cmd /c ver | Microsoft Windows … — same story, the kernel lives in WSL 2 |
| Linux | uname -sr | Linux 6.8.0-… — the very kernel your containers just reported |
So what's actually "Ubuntu" about an Ubuntu
container? Only the files — P2's tree: its own /etc, its own
apt, its own libc version:
docker run --rm ubuntu:24.04 head -2 /etc/os-release
docker run --rm alpine head -2 /etc/os-release
PRETTY_NAME="Ubuntu 24.04.4 LTS"
NAME="Ubuntu"
NAME="Alpine Linux"
ID=alpine
A Linux distribution is a kernel plus a userland — and in containers, the kernel is factored out and shared.² That's why an "operating system" image can be 8 MB (Alpine): it ships no kernel at all, just user-space files.
Why this matters for Docker & Kubernetes
Small images, instant starts, and dense packing all fall out of kernel-sharing — but so does the blast radius: a kernel panic, or a kernel that's too old for your app's syscalls, affects every container on the node. When Lesson 9's multi-arch section said images are per-CPU-architecture, this is why: user-space binaries must speak the shared kernel's architecture.
3.Lab B — strace, the x-ray machine
Bare Ubuntu can't show you syscalls — try it and meet a P4 exit-127-style error at the docker level:
docker run --rm ubuntu:24.04 strace ls
docker: Error response from daemon: … exec: "strace": executable file not found in $PATH: unknown.
Hence the oslab image you built. Enter it and x-ray a command you've used
since P1:³
docker run -it --rm oslab bash
strace -c ls / # -c: run it, then tally every syscall it made
% time seconds usecs/call calls errors syscall
------ ----------- ----------- --------- --------- ----------------
24.00 0.000360 25 14 mmap
20.80 0.000312 39 8 close
8.40 0.000126 18 7 fstat
5.73 0.000086 14 6 openat
4.53 0.000068 34 2 getdents64
4.07 0.000061 12 5 read
1.20 0.000018 18 1 write
0.00 0.000000 0 1 execve
Little ls crossed the wall ~60 times: execve to become
ls at all (P4's exec!), openat to reach the directory,
getdents64 to read its entries, one write to print — P3's
stdout, gate-level. Now narrow the x-ray to just the file-opens:
strace -e trace=openat cat /etc/hostname
openat(AT_FDCWD, "/etc/ld.so.cache", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/lib/aarch64-linux-gnu/libc.so.6", O_RDONLY|O_CLOEXEC) = 3
openat(AT_FDCWD, "/etc/hostname", O_RDONLY) = 3
01660e73a23e
+++ exited with 0 +++
Read it like a story: cat loads its libc, opens your file, and the kernel
answers each request with a small number — = 3 — a file
descriptor, the very same numbers P3 taught you as 0, 1, 2. Then the hostname
prints (it's the container ID — Lesson 1's UTS fact), and the process exits 0 — P4's exit
code, reported by strace itself.
Who's allowed to x-ray? (a Lesson 11 preview)
strace works via the ptrace syscall — spying on another process is
security-sensitive. Docker filters every container's syscalls through a
seccomp profile (a syscall firewall in the kernel); ptrace has been on
the default allow-list since Docker 19.03, which is why this worked with no extra
flags.⁴ On
older engines or hardened setups you'd add --cap-add SYS_PTRACE. Lesson 11
turns these dials the other way — removing capabilities.
4.Watch one command cross the wall
Here is that cat /etc/hostname trace again — as a film. Every prerequisite
lesson you've done is one of the frames:
- cat asks libc to open the file; libc prepares the raw openat syscall — the request approaches the gate.
- The CPU itself switches from user mode to kernel mode. This is the only legal way across the wall.
- Inside, the kernel walks the path and runs P2's exact permission algorithm — owner? group? other? — before agreeing.
- The kernel hands back a small number, fd 3 — P3's file descriptors were these all along.
- read(3, …) crosses again: the kernel copies the file's bytes into cat's memory.
- write(1, …) sends them out fd 1 — stdout — which your terminal is holding, exactly as P3 drew it.
- exit_group(0): the process ends and the shell collects P4's exit code. Every command you've ever run did all of this.
5.Lab C — /proc: the kernel, live, as files
P2 promised "everything is a file" gets weird and wonderful. Here it is:
/proc is a fake filesystem the kernel synthesizes on the fly — its live
internal state, dressed up as readable files.⁵
One numbered directory per process:
ls /proc | head -14
tr "\0" " " < /proc/1/cmdline; echo # what is PID 1 here, really?
ls -l /proc/self # /proc/self = "whoever is asking"
head -3 /proc/meminfo
1
7
8
buddyinfo
bus
cgroups
cmdline
consoles
cpuinfo
…
bash ← PID 1 is the bash you asked docker to run
lrwxrwxrwx 1 root root 0 Jul 17 03:16 /proc/self -> 9 ← points at the ls that asked!
MemTotal: 6081840 kB
MemFree: 249840 kB
MemAvailable: 2184212 kB
Three things worth savoring. The numbered dirs are the process list — and in
this container there are only a couple, P4's tiny tree. /proc/self is a
symlink that points at whichever process reads it. And MemTotal says
~6 GB — on Docker Desktop that's not your machine's RAM, it's the VM's allowance
(a Linux host would show its real total); the kernel answering is the one your
containers share.
Now the punchline: ps is not magic — it just reads
/proc. X-ray it:
strace -e trace=openat ps aux 2>&1 | grep "/proc" | head -8
openat(AT_FDCWD, "/proc/self/stat", O_RDONLY) = 3
openat(AT_FDCWD, "/proc/uptime", O_RDONLY) = 3
openat(AT_FDCWD, "/proc/12/status", O_RDONLY) = 3
openat(AT_FDCWD, "/proc/12/stat", O_RDONLY) = 3
openat(AT_FDCWD, "/proc", O_RDONLY|O_NONBLOCK|O_CLOEXEC|O_DIRECTORY) = 3
openat(AT_FDCWD, "/proc/uptime", O_RDONLY) = 4
openat(AT_FDCWD, "/proc/1/stat", O_RDONLY) = 4
openat(AT_FDCWD, "/proc/1/status", O_RDONLY) = 4
Every P4 tool decomposes this way: ps reads /proc,
docker stats reads cgroup files, ss reads
/proc/net. Tools are just polite formatting over kernel files.
/proc/PID/environ — why env vars aren't secrets
A process's environment (P4) is readable at /proc/PID/environ:
docker run --rm -e DB_PASSWORD=hunter2 ubuntu:24.04 sh -c 'tr "\0" "\n" < /proc/1/environ'
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
HOSTNAME=94dd9c883ef9
DB_PASSWORD=hunter2
HOME=/root
Anything that can read that file has your password — the kernel-level reason Lesson 4
said docker inspect shows env in plain text, and Lesson 11 moved secrets to
files.
6.Lab D — namespaces, by hand
Lesson 10 showed namespaces through Docker's flags. Here's the raw material. Which namespaces does your process live in? The kernel lists your memberships — as files, of course:
ls -l /proc/$$/ns # $$ = this shell's PID (P4)
lrwxrwxrwx 1 root root 0 Jul 17 03:16 cgroup -> cgroup:[4026532894]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 ipc -> ipc:[4026532815]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 mnt -> mnt:[4026532813]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 net -> net:[4026532818]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 pid -> pid:[4026532816]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 user -> user:[4026531837]
lrwxrwxrwx 1 root root 0 Jul 17 03:16 uts -> uts:[4026532814]
Each line is a room this process is locked into; the number identifies the room. Processes with the same number share that view of the world.⁶
| namespace | isolates | you met it in |
|---|---|---|
pid | which processes exist | Lesson 10's PID 1994-vs-PID 1 |
net | interfaces, IPs, ports | why -p exists (Lesson 4, P6 next) |
mnt | the filesystem view | image layers become / (Lessons 2–3) |
uts | hostname | container ID in the prompt (Lesson 1) |
ipc | shared-memory channels | (background — rarely touched by hand) |
user | uid/gid mapping | note its low number above: shared with the host — Docker doesn't unshare it by default, Lesson 11's rootless story |
Now build one. unshare is a small wrapper around the syscall of the same
name: "put my next command in new rooms".
docker run -it --rm --privileged ubuntu:24.04 bash
unshare --pid --fork --mount-proc bash # new PID namespace + fresh /proc
ps aux
exit # leave the namespace…
exit # …then the container
USER PID %CPU %MEM VSZ RSS TTY STAT START TIME COMMAND
root 1 0.0 0.0 4032 1272 ? S 03:16 0:00 bash
root 2 0.0 0.0 7628 3584 ? R 03:16 0:00 ps aux
Look at that ps: your bash is PID 1, and the world
contains two processes. That's Lesson 1's "sandboxed process" punchline and Lesson 10's
Lab A — reproduced with one command and zero Docker. You just built one-seventh
of a container. Docker's actual job is assembling all the rooms at once — plus cgroups,
plus the image's mnt-namespace filesystem, plus a nice API. "An isolated, restricted
process", nothing more mystical than that.²
--privileged is a learning-lab flag
Creating namespaces inside a container needs powers Docker rightly strips by default
(CAP_SYS_ADMIN among them⁷).
--privileged hands the container essentially the Docker host's full power —
fine for a throwaway experiment on your own machine, a serious finding anywhere near
production. Lesson 11 is the antidote chapter.
7.cgroups — the other half, briefly
Namespaces limit what a process can see. The kernel's other lever, cgroups, limits what it can use — memory, CPU, I/O. Same design language: the interface is files.⁸
docker run --rm --memory=64m ubuntu:24.04 cat /sys/fs/cgroup/memory.max
67108864
64 MiB in bytes. Docker's --memory flag did nothing clever — it wrote
that number into a cgroup-v2 file, and the kernel enforces it. When a process pushes past
the ceiling, the kernel's OOM killer ends it with SIGKILL — which P4
taught you to read in the exit code: 128 + 9 = 137. Lesson 10 stages
that whole crime scene deliberately; now you know both halves of its machinery.
You can now
- Draw the wall: user space, kernel space, and syscalls as the only gates — and x-ray
any command's gate traffic with
strace. - Explain why every container on a machine reports the same
uname -r, and what an image actually ships (userland files, no kernel). - Read the kernel's live state in
/proc— including whypsis just a file reader and env vars are never secrets. - Build a PID namespace by hand with
unshareand explain a container as an isolated (namespaces) + restricted (cgroups) process.
8.Check yourself
A system call is best described as…
- a library call inside libc
- the doorway into the kernel
- a process the kernel starts
- an interrupt the shell handles
Namespaces vs cgroups — which split is right?
- namespaces limit seeing; cgroups limit using
- namespaces limit using; cgroups limit seeing
- namespaces isolate files; cgroups isolate networks
- namespaces restrict memory; cgroups restrict processes
An Ubuntu container and an Alpine container on your machine print the
same uname -r. Why?
- Docker copies the kernel into images
- all containers share one host kernel
- uname is cached by the daemon
- both images ship identical kernel builds
What is /proc, and name one thing you'd read from it
while debugging.
A fake filesystem the kernel synthesizes: its live state exposed as files, one numbered directory per process. Debug reads: /proc/1/cmdline (what PID 1 really is), /proc/PID/environ (a process's env — why secrets leak), /proc/meminfo (the VM's real memory), /proc/PID/ns (namespace memberships).
You ran unshare --pid --fork --mount-proc bash and then
ps aux. What did ps show, and why?
Two processes, with bash as PID 1: unshare created a new PID namespace (bash renumbered to 1, everything else invisible) and mounted a fresh /proc so ps — which just reads /proc — sees only the new namespace's processes. That's the core of what Docker does when it starts a container.
9.Go deeper
Primary source: Ivan Velichko's Learning Containers From The Bottom Up — the "isolated + restricted process" mental model this lesson builds toward, with hands-on follow-ups on iximiuz labs. For strace joy, Julia Evans' strace posts & zine; for the dry truth, man7's namespaces(7) and proc(5).
Next: the net namespace got one row in a table — it deserves its own
lesson. P6: Ports, DNS & Packages opens
the network side: sockets, the 127.0.0.1-vs-0.0.0.0 trap behind every "why can't I reach
my container", and where apt actually puts things. And Lesson 10, when
you reach it later in the course, will read like a friendly rerun.
Stuck? Curious?
Bring questions to class, or open an issue on the course repo — include the command you ran and the output you got. The quizzes above are for self-checking: commit to an answer before revealing it, and re-try anything you missed tomorrow.
Sources
- man7 — syscalls(2) (what syscalls are; the full list)
- Ivan Velichko — Learning Containers From The Bottom Up (kernel+userland; isolated & restricted process)
- Julia Evans — strace posts & zine (strace technique)
- Docker Docs — seccomp profiles (syscall filtering; ptrace allowed since 19.03)
- man7 — proc(5) (/proc structure, environ, ns)
- man7 — namespaces(7) (namespace kinds and /proc/PID/ns)
- man7 — capabilities(7) (CAP_SYS_ADMIN and friends)
- man7 — cgroups(7) (cgroup v2 file interface)