Slow Jobs, Killer Containers: Laravel Queue Timeouts and Crashes in Docker
A job that sends an email takes half a second. A job that calls a generative model takes as long as the provider feels like, and sometimes a lot longer. Run that job in a Docker container and it has three ways to die: th
A job that sends an email takes half a second. A job that calls a generative model takes as long as the provider feels like, and sometimes a lot longer. Run that job in a Docker container and it has three ways to die: the worker timeout fires, the kernel kills it for memory, or a deploy pulls the container out from under it. Until this month Laravel treated those three deaths differently and didn't tell you. Laravel 13.33 and 13.34 change two of them. Here's which two, and what's still on you.
One job, three ways to die
Picture a job that sends an image to a model and waits for the answer, running under php artisan queue:work inside a container. In production it can end like this:
-
Worker timeout. The job exceeds
--timeout(or the job's$timeout). Laravel usespcntlandSIGALRM; when the alarm fires, the worker has always killed itself withSIGKILL. The timeout is recorded against the job (which fails immediately if it has$failOnTimeout) and the process is gone. -
Container OOM. The process goes over the container's memory limit (
mem_limitordeploy.resources.limits.memoryin Compose). The kernel sendsSIGKILL: no exception, nofailed()callback, no application log. Ifphpis PID 1, the container exits with 137. -
Deploy.
docker compose uprecreates the container. It sendsSIGTERM, the worker tries to finish its current job, and afterstop_grace_periodβ ten seconds by default β comesSIGKILL. For a job waiting on a model, ten seconds isn't much.
In the last two cases the job stays reserved until retry_after expires, then becomes available again and another worker picks it up. As far as the framework is concerned nothing happened: no exception was thrown, so no exception was counted.
The gap: maxExceptions can't see crashes
With $tries the damage is bounded: every time the job is popped the attempt counter goes up, and eventually it fails. But if you call rate-limited external APIs you often can't use $tries, because the RateLimited middleware releases the job and burns an attempt even when the job never actually ran. So the usual setup becomes $tries = 0, a $maxExceptions cap, and retryUntil() as a safety net.
The problem is that maxExceptions only counts real exceptions. A job that pushes the worker into OOM, gets killed, returns to the queue and pushes the next worker into OOM increments nothing. It loops until retryUntil() runs out, burning CPU and β if it calls a paid model β billed requests. The author of PR #61737 describes hitting exactly this with a worker being OOM-killed on Laravel Cloud.
The fix is opt-in, per job:
#[MaxExceptions(3), CountCrashesAsExceptions]
class GeneratePreview implements ShouldQueue
{
public $tries = 0;
public function retryUntil(): DateTime
{
return now()->addMinutes(30);
}
}
The mechanism is simple and honest. When the worker picks up the job it writes a marker to the cache, and deletes it when the attempt ends. If the marker is still there on the next attempt, the previous one died badly, and that death counts as one exception. It costs two cache calls per attempt, and it needs a cache shared between workers: with the array driver, or file inside separate containers, the marker dies with the container. It shipped in 13.34.0; instead of the attribute you can also set public $countCrashesAsExceptions = true;.
A timeout that doesn't kill the worker
PR #61591, shipped in 13.33, adds a static flag:
// AppServiceProvider::boot()
Worker::$killOnTimeout = false;
With it off, the worker doesn't kill itself when the timeout fires: it throws a TimeoutExceededException inside the job. The job gets a chance to clean up, close what it opened, even catch the exception and decide what to do. The worker survives and moves on to the next job without re-bootstrapping the framework and reopening every connection.
I'd be careful with it, for two reasons spelled out in the discussion on PR #61622. First, an exception thrown from a signal handler only propagates once PHP gets back to executing instructions; a job stuck inside a blocking system call never gets there. Second, if the exception does unwind, the worker carries on with whatever state the interrupted job left behind. SIGKILL is brutal, but it guarantees that state dies with the process. And a broad try/catch (\Throwable) inside handle() will swallow the exception and make the timeout disappear altogether β which is why the default stays true.
Line up your timeouts, inside out
Neither flag replaces the thing that actually matters: a strict ordering of timeouts, from the innermost to the outermost. In years of running Laravel queues this is the part I've seen get wrong most often, and rarely out of ignorance β the four numbers just live in four different files and nobody reads them side by side.
-
The HTTP timeout on the model call comes first.
Http::timeout(60)or your SDK's equivalent. It's the only one that turns waiting into a normal, catchable, counted exception. -
The job's
$timeout, a few seconds higher. It only fires if the first one wasn't enough. -
retry_afteron the queue connection, above the job timeout. If it's lower, a second worker picks the job up while the first is still running it. The docs say so; with a paid model it means paying twice for the same answer. -
stop_grace_periodon the Compose worker service, above the job timeout, so a deploy waits for the running job instead of killing it.
services:
worker:
image: app:latest
command: php artisan queue:work --timeout=90 --memory=256
stop_grace_period: 120s
deploy:
resources:
limits:
memory: 512M
A note on --memory: it stops the worker cleanly once memory goes over the threshold, but it checks between jobs. It won't save you from a single job that balloons past the container limit while decoding an image. For that you need a container limit with headroom, and β now β CountCrashesAsExceptions so you don't end up in a loop.
Your payload is a file, even if you don't call it one
One last point, from Miraviso, the SaaS I'm building for hair salons. There, the haircut preview is generated by Gemini on an EU server, with consent, and is never written to disk. That stack is FastAPI, not Laravel, but the rule translates word for word: put an image in a job payload and you've written it to your queue backend β Redis with persistence, a jobs table, and on failure failed_jobs. Retries read it back from there. I wrote about these undeclared copies in a piece on sensitive data in logs; queues follow the same logic. If the promise is "never on disk", an async job is the wrong home for that data, and automatic retries need rethinking accordingly.
The rest is hygiene: a PR that counts crashes, a flag that makes timeouts manageable, and four numbers in the right order. Of those, the four numbers are the only thing no upgrade will ever do for you.
Originally published at gabrielepieretti.dev. I write about Laravel, privacy-by-design and building a vertical B2B SaaS solo β more here.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.