Fork Once, Work Many: Why I Left Parallel::ForkManager Behind
Loading a GFS forecast hour used to mean forking a new process for every Param/Level record—roughly a thousand short-lived kids per forecast hour. Wx_GFS_DBLoad.pl now forks a small pool of long-lived workers once per forecast hour and feeds them tasks over pipes, cutting process churn while keeping concurrency. On ColoM1 (our colocated production server) the winning setup is a four-worker pipe into ClickHouse: less fork overhead, batched MariaDB bookkeeping, and about an 18% shorter load wall than the old fork-per-record path at the same worker count.
TL;DR
Wx_GFS_DBLoad.pl used to lean on Parallel::ForkManager and spin up a fresh child for every Param/Level record—about a thousand forks per forecast hour. We still stay in Perl, but for each forecast hour we now fork a small pool of long-lived workers once, feed them tasks over pipes, and let the parent handle MariaDB. On ColoM1 (our colocated production server), a four-worker pipe path into ClickHouse beats the old four-fork-per-record approach on wall time (~18% in our late-September A/B), with more of the CPU spent doing real work instead of process churn. Fewer forks isn’t always gentler—or faster.
The problem: concurrency with a cost
GFS DB load is embarrassingly parallel at the record level. For each forecast hour you wait on the Primary and Secondary GRIB pair, pull Param/Level work from the .idx, run wgrib2, and land fields in ClickHouse while MariaDB tracks status in Weather.GFS_Status.
Parallel::ForkManager made that easy: one child per record, bounded by --forks. It also meant roughly a thousand process start/exit cycles every forecast hour. Each kid paid fork tax, opened its world, ran wgrib2 → ClickHouse, and died. Concurrency was real. So was churn.
I wasn’t looking for a language rewrite. I wanted the same Perl codebase, the same job shape, and less overhead per task.
What changed: per-hour worker pool
As of the 2026-09-21 change, the model is: fork $forks workers once per forecast hour, then feed them.
- Parent waits for the Primary+Secondary pair, batch-marks Translating, builds tasks from the index, upserts Param/Level rows in chunks (parent only), disconnects MariaDB.
- Fork N long-lived children. Each gets a task pipe and a result pipe (length-prefixed Storable messages, autoflush).
- Round-robin the task list across workers; close the task pipes when the queue is empty so workers exit cleanly.
- Parent collects results (progress every 10%; optional
--stageaverages for wgrib2 / insert / compress). waitpidthe pool; reconnect MariaDB; batch status updates. Any child failure can mark the Full rowError: DBLoad.
Workers never touch DBI. They only run gfsProcessOneTask—extraction and ClickHouse load. The default path pipes wgrib2 stdout straight into ClickHouse (no intermediate CSV on the prod path). --csv / --gzip / --zstd are still there when you want files for debugging.
--forks still sets pool size. If you omit it, cpuForks() picks from load and core count (prod floor 4). Prod settled on four workers in cron.
Wall time vs system load (what we actually measured)
These are internal A/B timings from specific late-September 2026 cycles on prod (and a few MikeM1—the development laptop—debug twins)—Apple Silicon boxes with our MariaDB + ClickHouse inventory. Not a universal ForkManager-vs-pool benchmark.
On prod (pipe path unless noted):
Setup | Rough mean per forecast hour | Notes |
|---|---|---|
CSV path, 3 forks | ~21.5–21.6 min/h | Older file-based baseline |
Pipe, 2 workers | ~25.8 min/h | Lighter-looking load, worse wall |
Pipe, 3 workers | ~19.8 min/h | ~8% better than CSV 3-fork |
Pipe, 4 forks (ForkManager-era, fork-per-record) | ~17.7–18.0 min/h | Solid, but still pay fork-per-task |
Pipe, 4 worker pool | ~14.8 min/h (about 14.5–15.3) | Same concurrency, less churn |
One clean 18z pool run: wall about 3:08, load-only closer to 2:58, versus roughly 3:36 on the prior four-fork-per-record style. P-core residency climbed (more time actually busy ~3 GHz), SSD throughput up a bit, zero fails on that pass.
So the winning story isn’t “pool is slower but kinder.” At the same worker count, the pool was faster on wall clock because we stopped paying fork-per-record tax and stopped chatting with MariaDB from every child. We didn’t invent magic cores; we wasted fewer of them on process theater.
Dev remains a debug twin (similar file path, fewer forecast hours). Under --stage, typical piece times there looked like wgrib2 ~2.1–2.5s and insert ~1.2–1.5s—useful for tuning, not the production headline. Several command line options allow testing and tweaking, while the underlying code remains identical.
Tradeoffs worth saying out loud
- Under-forking hurts. Two workers looked “gentler” and lost hard on wall time. Throughput matters when a forecast hour has a thousand tasks.
- Pool gentleness is mostly about churn, not about soft-pedaling ClickHouse. Fork once per forecast hour; MariaDB stays in the parent; workers stay focused.
- Progress logs can fib for a minute. With 10% steps and second-resolution timestamps, the first burst of completions can look simultaneous. That’s queue dynamics, not stuck work.
- OPTIMIZE is its own weather system. Duration swings with merges and locks. Don’t treat one short OPTIMIZE as proof the loader got faster.
Other September iterations: keepers and compost
Kept:
- Wait for Primary+Secondary before translating a forecast hour (no half-pairs).
- Batch MariaDB upserts and status updates; fail the Full row if any child fails.
- Pipe
wgrib2→ ClickHouse on prod (drop CSV files on the prod path); later made pipe the default everywhere, with file flags opt-in. --stagefor profiling, default off once we trusted the timings.- Stay in Perl with a worker pool instead of rewriting the world.
OPTIMIZE … FINAL,optimize_throw_if_noop, and a short retry/Slack path after silent no-ops lied to us.
Composted:
- macOS
taskpolicy/ QoS A/B hunting Performance cores — no win on dev or prod; ripped out. - “Fewer forks must be better” — the 2-worker runs voted no.
- Trusting non-FINAL OPTIMIZE during concurrent inserts (Ok with MergedRows=0 is not a victory lap).
The backup tree tells the same story if you like archaeology: bak-pool-*, bak-pipech-*, bak-batchdb-*, bak-notaskpolicy-*.
Takeaway
Parallel::ForkManager is a great hammer when tasks are coarse and fork cost is noise. When you’re forking a thousand times per forecast hour to run the same shape of wgrib2 → ClickHouse work, the hammer starts charging rent. A small, long-lived worker pool with piped tasks kept our concurrency, cut process and DB-connection churn, and, on prod at four workers, actually shortened the wall clock.
Same language. Same forecast-hour job. Less ceremony per record.
Next up in this series: more of what happens after the grids land. For now, the loader spends more time loading and less time auditioning for process theater.
Stay tuned, and happy forecasting!