Docker Deep Dive · Prerequisites · P4 of 7

Processes, Environments & Signals

Everything a container really is — a process — gets built here: PIDs and parent trees, the env-var table Docker injects into, and the two signals that decide whether your app shuts down gracefully or gets killed at t=10s.

⏱ ~25 min hands-on · repo chapter p4-processes-env-and-signals · cheatsheet: Linux & Shell

prereq P4 / 7

Before you start

cd docker-deep-dive
git pull
cd p4-processes-env-and-signals
docker run -it --name procbox -v "$(pwd)":/labs ubuntu:24.04 bash
docker run -it --name procbox -v "${PWD}:/labs" ubuntu:24.04 bash

That last command is your Linux lab for this lesson. -v "$(pwd)":/labs shares this chapter folder into the container at /labs (Lesson 5 explains the flag properly — for now: shared folder). Type exit any time; docker start -ai procbox resumes it.

1.A program is a file. A process is a life.

/usr/bin/sleep sitting on disk is a program. The moment you run it, the kernel creates a process: the program's code plus everything it needs to live — a PID (process id), a parent (PPID), its own memory, its three streams and any other open files (P3), a working directory (P1), a private copy of the environment (§3), and one slot for the exit code its parent will collect when it dies.¹

Inside procbox, start two background processes and take the census:

sleep 300 &
sleep 400 &
ps aux
USER       PID %CPU %MEM    VSZ   RSS TTY      STAT START   TIME COMMAND
root         1  0.0  0.0   4032  3016 ?        Ss   03:15   0:00 bash
root         8  0.0  0.0   2268   768 ?        S    03:15   0:00 sleep 300
root         9  0.0  0.0   2268   748 ?        S    03:15   0:00 sleep 400
root        10  0.0  0.0   7628  3204 ?        R    03:15   0:00 ps aux

Reading the columns that matter: PID — the id you'll kill; STAT — S sleeping, R running, (Z would be a zombie, §4); TIME — CPU actually used, which is why two "300-second" sleeps show 0:00. And the headline: your shell is PID 1. In a container, the first process gets the first number — the whole "container = a process" idea from Lesson 1, now visible from the inside.

Every process has exactly one parent — whoever spawned it. ps -ef adds the PPID column:

ps -ef --forest
UID        PID  PPID  C STIME TTY          TIME CMD
root         1     0  0 03:15 ?        00:00:00 bash
root         9     1  0 03:15 ?        00:00:00 sleep 300
root        10     1  0 03:15 ?        00:00:00 sleep 400
root        11     1  0 03:15 ?        00:00:00 ps -ef --forest

PPID 1 on every child: bash begat them all. Processes form a tree, and PID 1 is its root. You'll see this tree from both sides of a container wall — docker top (host view) vs docker exec … ps (inside view) — in Lesson 10, later in this course; P5 explains the namespace trick that makes both true at once.

2.Foreground, background, and what Ctrl-C really is

The shell runs one foreground process at a time — that's why your prompt disappears until a command finishes. The P1 keys were signals in disguise all along:

sleep 500 &
sleep 600 &
jobs
[1]-  Running                 sleep 500 &
[2]+  Running                 sleep 600 &

Try it: fg %1, then Ctrl-C it dead. Check jobs again — one left.

3.The environment: a table every process carries

Each process owns a private KEY=value table — its environment. At spawn, the child gets a copy of the parent's table; after that the copies are independent. Changes never flow back up, and never sideways into already-running children.³ The subtlety that bites everyone: a plain shell variable is not in that table until you export it. Prove all three cases in one go:

X=1 bash -c 'echo one-shot: ${X:-unset}'
X=2
bash -c 'echo plain: ${X:-unset}'
export X=3
bash -c 'echo exported: ${X:-unset}'
one-shot: 1
plain: unset
exported: 3

VAR=x cmd injects a value for that one child only; a bare assignment stays private to the shell; export puts it in the table children copy. (${X:-unset} means "X, or the word unset if it's empty" — P7 builds entrypoints out of that.)

your shell · PID 40 PATH=/usr/local/bin:/usr/bin… HOME=/root X=3 (exported ✓) name=local (no export — stays here) spawn = COPY exported rows only python app.py · PID 57 PATH=/usr/local/bin:/usr/bin… HOME=/root X=3 (its own copy now) ✗ child edits never flow back (or sideways)
The environment is copied at spawn, one direction only. docker run -e PORT=9000 (Lesson 4) does nothing more mysterious than write a row into this table before the container's PID 1 starts — which is exactly why 12-factor apps read config from it.⁶

One env var runs the whole show: PATH, the colon-separated list of directories the shell searches, left to right, when you type a bare command name (P1's "how does it find ls?", answered). Break it and watch:

PATH='' /usr/bin/ls /etc/hostname
PATH='' ls /etc/hostname
echo "after failure: exit=$?"
/etc/hostname
bash: line 1: ls: No such file or directory
after failure: exit=127

Full paths need no search; bare names die without one — and 127 ("command not found") joins your exit-code vocabulary. This is why Dockerfiles ENV PATH=/app/venv/bin:$PATH to put a virtualenv "first in line".

4.Signals: how processes are told things

A signal is a one-byte async message the kernel delivers to a process. There are ~64 (kill -l lists them); four run your life:²

signalnmeaningcatchable?
SIGINT2Ctrl-C — "stop what you're doing"yes
SIGTERM15"please exit" — default of kill and docker stopyes — graceful shutdown lives here
SIGKILL9kernel deletes the process. Not a request.never
SIGHUP1terminal closed; daemons reuse it as "reload config"yes

"Catchable" means the process may install a handler — code that runs when the signal lands (bash calls this trap, P7 writes one). A process that catches SIGTERM can flush buffers, close connections, then exit. SIGKILL skips the process entirely — the kernel just removes it, which is also why it can leave half-written files behind.

Death leaves a number. $? holds the last command's exit code: 0 is the only success; 126/127 are the shell's "found but not executable"/"not found"; and a process killed by signal N exits 128 + N.⁸ Watch both flavors:

sleep 500 &
kill $!
wait $!
echo "polite kill -> exit=$?"
sleep 500 &
kill -9 $!
wait $!
echo "kill -9   -> exit=$?"
polite kill -> exit=143
kill -9   -> exit=137

143 = 128+15 (TERM), 137 = 128+9 (KILL) — the exact number Lesson 10's OOM kill produced. From now on, exit 137 reads as "something SIGKILLed it": the OOM killer, docker stop's deadline, or an impatient human.

The PID 1 twist — why containers ignore polite requests

One rule changes everything for containers: for PID 1, the kernel discards any signal the process hasn't installed a handler for (ordinary processes die by default; PID 1 doesn't).⁷⁴ Your app is PID 1 in its container. No SIGTERM handler → docker stop's polite request evaporates. Prove it from a second terminal on your machine (leave procbox running):

docker run -d --name p4-naive -v "$(pwd)":/labs ubuntu:24.04 bash /labs/naive-worker.sh
docker exec p4-naive kill -TERM 1
docker exec p4-naive ps -ef
docker run -d --name p4-naive -v "${PWD}:/labs" ubuntu:24.04 bash /labs/naive-worker.sh
UID        PID  PPID  C STIME TTY          TIME CMD
root         1     0  0 03:16 ?        00:00:00 bash /labs/naive-worker.sh
root        25     1  0 03:16 ?        00:00:00 sleep 1
root        26     0  0 03:16 ?        00:00:00 ps -ef

…nothing. Still working. That's not a bug in kill — it's the PID-1 rule, and it's why the next section's animation is the single most useful mental model in container operations. (Related PID-1 duty, for honesty's sake: when a child process dies, its entry lingers as a zombie until the parent collects the exit code with wait(). PID 1 inherits every orphan, so a PID 1 that never wait()s slowly fills the process table — the reason serious images use tiny init shims like tini, mentioned again in Lesson 11.)

docker rm -f p4-naive

5.Anatomy of docker stop

dockerd docker stop web container "web" PID 1 · your app child: sleep 1 working 03:16:57 · working 03:16:58 · … SIGTERM (15) ⏱ t = 0 s path A — handler installed trap runs: "SIGTERM caught — flushing…" app closes connections, exits 0 Exited (0) — in 0.2 s ✓ 🛡 path B — no handler: kernel DISCARDS the signal for PID 1 ⏱ t = 10 s — grace over SIGKILL (9) no appeal — Exited (137) = 128 + 9 moral: PID 1 must handle SIGTERM — P7 scripts it; Kubernetes plays the same game (30 s grace)
  1. docker stop asks the daemon to deliver SIGTERM to the container's PID 1 — and starts a 10-second stopwatch.
  2. Path A: the app installed a handler. Cleanup runs, it exits 0, the container parks as Exited — all in a fraction of a second.
  3. Path B: no handler. An ordinary process would die by default, but for PID 1 the kernel simply discards the signal. The app doesn't even notice.
  4. So nothing happens… for the whole grace period (Kubernetes does the identical dance with terminationGracePeriodSeconds: 30).
  5. Deadline. The daemon sends SIGKILL — uncatchable even for PID 1. No cleanup, exit code 137.
  6. The moral of phase 0's most operational lesson: handle SIGTERM, or every deploy/restart of your app is a small crash.
Two endings, ten seconds apart. docker stop = SIGTERM, wait (default 10 s, -t to change), SIGKILL.⁵ Lesson 12 taught Harbor Log the path-A ending; here you've seen the mechanics that force the choice.

6.Lab — feel the ten seconds

The chapter ships both workers; race them. From the chapter folder on your machine (macOS / Linux / WSL below — the PowerShell spelling of the whole race follows):

docker run -d --name p4-naive -v "$(pwd)":/labs ubuntu:24.04 bash /labs/naive-worker.sh
time docker stop p4-naive
docker inspect --format 'ExitCode={{.State.ExitCode}}' p4-naive
docker stop p4-naive  0.06s user 0.03s system 0% cpu 10.229 total
ExitCode=137
docker run -d --name p4-graceful -v "$(pwd)":/labs ubuntu:24.04 bash /labs/graceful-worker.sh
time docker stop p4-graceful
docker logs --tail 3 p4-graceful
docker inspect --format 'ExitCode={{.State.ExitCode}}' p4-graceful
docker stop p4-graceful  0.07s user 0.04s system 51% cpu 0.209 total
working 03:16:59
working 03:17:00
SIGTERM caught — flushing and exiting cleanly
ExitCode=0
docker run -d --name p4-naive -v "${PWD}:/labs" ubuntu:24.04 bash /labs/naive-worker.sh
Measure-Command { docker stop p4-naive }
docker inspect --format 'ExitCode={{.State.ExitCode}}' p4-naive
docker run -d --name p4-graceful -v "${PWD}:/labs" ubuntu:24.04 bash /labs/graceful-worker.sh
Measure-Command { docker stop p4-graceful }
docker logs --tail 3 p4-graceful
docker inspect --format 'ExitCode={{.State.ExitCode}}' p4-graceful

10.229 s vs 0.209 s — a 50× difference, and the clean one also got to say goodbye (flush logs, close the database connection, finish the request). Open graceful-worker.sh — the whole trick is four lines of trap, and P7 turns it into the standard entrypoint pattern. Then tidy up:

docker exec -w /labs procbox bash check.sh
docker rm p4-naive p4-graceful
docker rm -f procbox

You can now

7.Check yourself

A container exited with code 137. What happened?

  1. the process finished its work normally
  2. the kernel killed it with SIGKILL
  3. the process caught SIGTERM and exited
  4. the shell could not find it

In bash you run X=1 (no export), then bash -c 'echo ${X:-unset}'. What prints?

  1. 1, because children copy every variable
  2. 1, but only inside login shells
  3. unset, since only exported variables inherit
  4. an error, because X is undefined

A process installs handlers for every signal it can. SIGKILL arrives. What happens?

  1. its handler runs and ignores it
  2. the kernel terminates it regardless anyway
  3. the signal queues until handlers finish
  4. the shell converts it to SIGTERM

SIGTERM vs SIGKILL — one breath each.

SIGTERM (15) is a catchable request: a handler may clean up and exit (unhandled, a normal process dies with exit 143). SIGKILL (9) is not a request: the kernel removes the process immediately, no handler, no cleanup — exit 137.

Why does a trap-less PID 1 survive docker stop's SIGTERM — and what happens at t=10s?

For PID 1 the kernel discards signals that have no installed handler (ordinary processes would die by default), so the SIGTERM evaporates. When the grace period expires the daemon sends SIGKILL, which even PID 1 cannot ignore — the container exits 137.

8.Go deeper

Primary source: The Linux Command Line ch. 10 ("Processes") — free at linuxcommand.org — plus the signal(7) man page for the full signal table. Julia Evans' "Should you be scared of signals?" is the friendliest second pass.

Next: you now know a container is a process with a PID 1, an env table, and a signal mailbox. P5 opens the trapdoor: what the kernel actually is, how strace shows every favor a process asks of it, and how namespaces make your PID 1 believe its own story.

Stuck? Curious?

Bring questions to class, or open an issue on the course repo — include the command you ran and the output you got. The quizzes above are for self-checking: commit to an answer before revealing it, and re-try anything you missed tomorrow.

Sources

  1. Shotts, The Linux Command Line, ch. 10 — Processes
  2. man7 — signal(7) (signal table, default actions, catchability)
  3. man7 — environ(7) (per-process environment, inheritance)
  4. Docker Docs — container run reference (PID 1 signal note)
  5. Docker Docs — docker container stop (SIGTERM, grace period, -t)
  6. The Twelve-Factor App — Config & Disposability
  7. man7 — pid_namespaces(7) (the PID-1 signal rule)
  8. GNU Bash manual — Exit Status (127, 128+N)