Three Years of Docker, Then Podman: How We Cut Memory 40% and Fixed Our Architecture

·Platform Decision·9 min read

Translated from the original Korean post. 한국어 원문 보기 →

2 AM, and the Cluster Fell Over in 60 Seconds

Sixty seconds. That's how long it took for three nodes to go down. Forty-two microservices across six clusters stopped at once, and the runtimes were cycling through restarts like a broken elevator. I opened my laptop in the dark, pulled up a process monitor, and stared at something I didn't believe.

A single dockerd process was holding 5.8 gigabytes of resident memory. On a host running nothing but plain Go services. But the number wasn't the worst part. The worst part was the log line sitting right above it.

"dockerd memory pressure detected"

I'd seen that warning before. Several times. Every time I told myself it was probably fine. Engineers lie to themselves the way drivers pretend not to hear the noise coming from under the hood. I was no different.

The Warnings I Kept Ignoring

Here's the embarrassing part. I'd already read an article about Podman back in 2021. I opened the tab, skimmed the first paragraph, closed it. Switching tooling felt risky, and Docker was already in my hands.

Looking back, the most expensive line item in infrastructure engineering was familiarity. It never shows up on an invoice. That's what makes it dangerous. It doesn't disappear just because nobody billed you for it — it accumulates somewhere, and then the whole balance comes due on the worst possible night. Same thing happened when I ran internet banking at a financial institution. Deferred decisions don't evaporate. They come back with interest, on the day traffic spikes.

The next morning I sent Priya, my co-lead, the screenshots from the night before. No explanation. She looked at the screen for a while, then said:

"So we've been paying money to feed the demon that eats our lunch."

She wasn't wrong. Priya is almost never wrong, which is occasionally inconvenient, but this time she was exactly right. That one line was the actual trigger for the migration.

Why Podman Is Fundamentally Different

To understand the gap between Docker and Podman, you have to start with the architecture, not the commands. Docker routes everything through a central daemon. One big background process governs the entire container world. Convenient. Also: when that process wobbles, everything sitting on top of it wobbles too. Restart dockerd? Every container restarts with it.

This is a single point of failure, reproduced faithfully at the operational layer. Running internet banking, the picture I feared most was "all traffic passes through one component" — and dockerd was sitting in exactly that seat. On the architecture diagram it's one small box. When that box dies, every line drawn above it goes dead. You don't see it in the drawing. You see it at 2 AM.

Podman starts from a different premise. Each container is just a child process of whatever launched it. No heavy daemon in the middle. One container dies, and nothing else on the host notices.

That distinction sounds academic until you're staring at six gigabytes of resident memory at the worst possible moment. Then it stops sounding academic.

The second thing that surprised me was rootless mode. Podman was designed rootless from day one. Running containers without root wasn't optional for us, and that single property cut our security review from two weeks to three days. Whether the daemon runs as root or not — for a security team, that difference is bigger than you'd think. Remove one privilege boundary and you remove that much of the threat model they have to reason about.

The Migration Was Simpler Than Expected

There was almost no barrier to entry. Same Dockerfile syntax, same OCI images, same registries. Even the commands are nearly identical.

podman build -t api:v2 .
podman run -d -p 8080:8080 api:v2
podman ps
alias docker=podman

We hammered on it in staging for two weeks looking for edge cases. Most of our scripts ran untouched. You can tell when a tool was built with compatibility in mind — it shows up in exactly these moments. Every command reflects how carefully somebody worked to avoid breaking existing users' muscle memory.

How We Actually Rolled It Out

We didn't flip the switch all at once. If working as an ops PM taught me one thing, it's that "everything at once" migrations end with somebody rolling back at 3 AM. So we picked the safest possible ordering.

The most boring service went first. An internal metrics shipper that could be dead for hours before anyone noticed. Ten days on Podman. Then read-only APIs. Then the worker pools. Smallest blast radius first, lightest traffic first. A migration is a risk-allocation exercise before it's an engineering one. You have to break things where it doesn't hurt, so you know what to watch for when you move the parts that do.

The real obstacle was Docker Compose. Podman has its own Compose implementation, but some of the networking flags we depended on hadn't caught up yet.

We ended up rewriting those files as Podman Quadlets — systemd unit files built for containers.

[Unit]
Description=API service

[Container]
Image=registry.local/api:v2
PublishPort=8080:8080
Environment=ENV=prod

[Install]
WantedBy=default.target

Now systemctl start api brings the service up. systemd handles restarts. Logs go to journald. No separate daemon supervising any of it. Container management moved inside the OS's standard process management, and that was the whole point. We took a layer we'd been building and maintaining ourselves and handed it to a system the OS has been refining for decades. Clean.

What the Numbers Said

After eight weeks, one full cluster was moved and I went looking at the metrics. Honestly, I'd have called 15% a win. The actual numbers blew past that.

Metric Docker Podman
RAM (avg) 4.9GB 2.9GB
RAM (peak) 7.2GB 3.4GB
Startup time 1.8s 1.1s
Failed pulls 0.7% 0.2%

I showed Priya the table the next morning. She glanced at it and said, "That's four instances." That was the whole comment. We shut down four EC2 instances. All from removing a middleman nobody had ever deliberately chosen.

What I find interesting is that none of this came from new investment. We didn't buy hardware. We didn't move to pricier instances. We peeled off the overhead the daemon had been eating, and the space it vacated came back as money. On consulting engagements the cost-cutting conversation always drifts toward "what should we buy," when the big chunk is usually hiding in "what we already bought and have been leaking." This was that, exactly. And the yearly savings bought back engineering capacity that had been quietly draining the whole time.

The Architecture Change, Drawn Out

Here's the whiteboard sketch I drew for the team.

Before (Docker)

engineer -> Docker CLI -> dockerd (the hungry daemon)
                              |
            +-----------------+-----------------+
            |                 |                 |
        container         container         container

After (Podman + systemd)

engineer -> podman -> container (owned by systemd)
engineer -> podman -> container (owned by systemd)
engineer -> podman -> container (owned by systemd)

The middleman is gone. Each container runs fully independently, and the blast path for any failure got that much shorter. Once the lines stop converging on a center point, there's simply no channel left for a failure to spread through.

The Limitations, Stated Honestly

Podman is not a perfect tool. If you're considering the migration, go in knowing these.

First, compatibility. Some third-party systems still assume the Docker socket is sitting right where it always was. To keep our CI runners from erroring out, we had to expose the Podman socket in compatibility mode. So we didn't actually delete Docker — we slid in a thin shim that convinces things Docker is still there. Anyone who's stripped out legacy has seen this shape before. The dependency you thought you removed is alive behind a thin adapter.

Second, networking performance. Rootless networking on slirp4netns is noticeably slower than bridge networking. On the worker boxes that need maximum throughput, we still run Podman as root. That's a deliberate compromise between security and throughput. We didn't force the same answer onto every host.

Third, if you're on Docker Swarm, Podman is the wrong choice. No need to soften that one.

These are real tradeoffs you have to absorb. Recommending the switch while hiding them wouldn't be honest.

What Three Years of Running This Taught Me

Three years of running Docker, then moving to Podman, and the biggest thing I'm left with is this: a tool's underlying philosophy gets billed back to you in production. Docker's centralized daemon was comfortable early and became a single point of failure as we scaled. Early convenience and later risk have to be evaluated separately, and at adoption time you can't see the second one. Convenience registers on day one. Risk gets billed at 2 AM a few years later.

Podman's daemonless model felt strange at first, but with containers running independently, the whole system became more predictable. And the security upside from rootless mode honestly exceeded what I expected.

40% less memory is a real number, but the predictability matters more. Nothing wakes me up at 2 AM because of dockerd anymore. For an operator, a night of decent sleep is a more honest metric than any benchmark.

Podman isn't the right answer everywhere. But if Docker's resource footprint or stability keeps nagging at you, it's worth a serious look.

Choosing technology comes down to picking the tool that fits your situation. The step before that — pulling your foot out of the familiarity trap — takes more nerve than I expected. I found mine at 2 AM.

Was this post helpful?

One click helps me write the next one

#Docker#Podman#Containers#Migration#Performance Optimization