Mastering Docker Multi-stage Builds
A single-stage Dockerfile ships everything it needed to build. The compiler, the package manager, the dev dependencies, the source, the intermediate object files and every apt package you installed to make one of those work are all still there in the final image. That is how a Node service whose output is eight megabytes of JavaScript ends up as a 1.2 GB image.
Multi-stage builds fix this by letting one Dockerfile define several independent images and copy files between them. Only the last stage becomes the result; everything else is scaffolding that is thrown away.
The shape of it
# syntax=docker/dockerfile:1
FROM node:22-bookworm AS build
WORKDIR /app
COPY package.json pnpm-lock.yaml ./
RUN corepack enable && pnpm install --frozen-lockfile
COPY . .
RUN pnpm build && pnpm prune --prod
FROM node:22-bookworm-slim AS runtime
WORKDIR /app
ENV NODE_ENV=production
COPY --from=build /app/node_modules ./node_modules
COPY --from=build /app/dist ./dist
USER node
CMD ["node", "dist/server.js"]
Two stages, named with AS. COPY --from=build reaches into the earlier stage’s
filesystem and takes only what is named. The compiler toolchain, the dev
dependencies and the source tree never enter the runtime image — so they cannot
be exfiltrated from a compromised container, and they cannot show up in a
vulnerability scan you then have to triage.
The gain is largest for compiled languages. A Go binary copied into scratch or
gcr.io/distroless/static produces an image of a few megabytes containing
exactly one executable and no shell for an attacker to find.
The layer cache is the other half
Every instruction produces a layer. Docker reuses a cached layer when the instruction and its inputs are unchanged — and critically, a cache miss invalidates every layer after it. So the ordering of instructions is a performance decision, not a style one.
That is why the dependency manifest is copied and installed before the source is copied. Dependencies change rarely; source changes constantly. Copy everything up front and every one-line edit reinstalls the entire dependency tree.
For COPY, cache validity is computed from the contents of the files copied.
For RUN, it is the literal command string — which is why RUN apt-get update
on its own line is a trap: the string never changes, the layer is reused for
months, and it hands a stale package index to the apt-get install below it.
Keep them in one RUN.
A .dockerignore file is not optional here. Without it, COPY . . pulls in
node_modules, .git and build output from your laptop, which both bloats the
build context and busts the cache whenever any of them changes.
BuildKit: cache mounts and parallelism
BuildKit is the default builder and it changes what is possible. The most
valuable feature is the cache mount, which gives a RUN step a persistent
directory that survives between builds and is not stored in the image:
RUN --mount=type=cache,target=/root/.local/share/pnpm/store \
pnpm install --frozen-lockfile
Now even a genuine cache miss on the lockfile does not re-download packages —
the package store is still on disk. The same trick applies to ~/.cargo,
~/.m2, Go’s module cache and apt’s archive directory. Secrets get a similar
treatment with --mount=type=secret, which exposes a file for the duration of
one instruction and leaves nothing in any layer, unlike an ARG or ENV that
is permanently recorded in the image history.
BuildKit also builds a directed graph of the stages rather than running them top to bottom. Stages that do not depend on each other build in parallel, and stages the final image never copies from are skipped entirely.
Patterns worth stealing
A test stage that is never shipped. Add FROM build AS test running the
suite, and invoke it in CI with docker build --target test .. Tests run in the
identical environment as the build, with no separate compose file.
A shared base stage. When several stages need the same system packages,
factor them into FROM node:22 AS base and have both build and test extend
it. One install, cached once.
Pin the base image. node:22-bookworm moves. For reproducible builds pin the
digest — node:22-bookworm@sha256:... — and update it deliberately.
Match the base to the artefact. -slim variants drop most of the userland.
Alpine is smaller still, but it uses musl instead of glibc, which changes DNS
resolution behaviour and can measurably slow allocation-heavy runtimes; verify
rather than assume. Distroless images have no shell at all — excellent for
production, awkward the first time you need to debug one.
Run as a non-root user. Containers default to root, and a root process inside a container is one misconfiguration away from being root-adjacent outside it.
Measuring
docker history IMAGE shows the size each layer contributed and is the fastest
way to find the 400 MB RUN you forgot about. docker build --progress=plain
shows which steps were CACHED and which re-executed, which is how you confirm
an ordering change actually helped rather than assuming it did.
Quick answers
- What is a Docker multi-stage build?
- A Dockerfile with more than one FROM, where earlier stages compile or install and a final stage copies only the finished artefacts. Build tools, compilers and source never reach the shipped image.
- Do multi-stage builds make images smaller?
- Substantially, because the final image contains only what you copy into it. A Node or Go build that ships without its toolchain and source is routinely an order of magnitude smaller than a single-stage equivalent.
- Can you copy from a specific stage?
- Yes — name stages with FROM image AS builder and then COPY --from=builder. You can also copy directly from an external image, which is a neat way to pull in a single binary without installing anything.
- Do multi-stage builds make builds faster?
- Only indirectly. They let independent stages be cached and, with BuildKit, built in parallel. The bigger speed win is still instruction ordering so dependency installs stay cached.
References
Related Discoveries
Lumi's weekly note
A short email when we publish something useful. No spam, unsubscribe anytime.