A queue gets work to a process. It does not, by itself, answer who may finish that work after something goes wrong.
Imagine a reporting job that takes a few minutes. A worker receives it and begins collecting data. Then it pauses long enough to miss its lease renewal. The system makes the job available again, and another worker claims it.
Eventually, the first worker wakes up. It still has the job in memory. It may even have a perfectly reasonable result. Now two attempts can try to publish the same report.
The useful question is not just whether the job will run again. It is: who is allowed to make the next change, and where is that permission enforced?
This is a hypothetical example, not an incident report. The timings and identifiers below illustrate the design rather than describe a particular system.
The failure is a handoff, not a disappearance
Suppose each claim has a generation number that increases when a new attempt takes ownership:
00s A claims the job, generation 41; lease ends at 30s.
10s A pauses and stops renewing.
30s A's lease expires. A has not necessarily stopped.
31s B claims the job, generation 42.
35s A resumes and submits a result with generation 41.
36s The result store rejects A's stale write.
40s B submits with generation 42; its write is accepted.The rejection at 36 seconds is the important part. A timeout made recovery possible. A guarded write made that recovery safe.
Ownership needs a deadline
A lease gives an attempt permission to work for a bounded period. The claim records an owner and an expiration time. Acquisition must be atomic: two workers cannot both win the same claim. Renewal must also check the current owner and refuse to revive an already expired claim.
Expiration makes abandoned work eligible for recovery. It does not stop the original process from running. The old worker might wake up after a pause and continue with assumptions that are no longer true.
That distinction is the foundation of the design: expiration changes permission, not process state.
Choose a lease duration and renewal interval with room for ordinary delays, but do not treat a generous timeout as a correctness guarantee. If a worker cannot confirm renewal before its known deadline, it should stop starting new side effects. That cooperative behavior helps; the result store still needs to defend itself against a delayed or paused worker.
A stale worker must lose its authority
For a Redis-backed claim, use an atomic acquisition and a unique ownership token. Release must verify that token; an unconditional deletion can remove a newer worker's claim. The Redis locking documentation explains the ownership checks, failover limitations, and clock assumptions involved. A key with a time-to-live is not a blanket guarantee of exclusive execution.
A random ownership token answers, "Is this still my claim?" An increasing generation answers, "Has a newer claim superseded mine?" They serve different purposes. A downstream resource that supports fencing can reject an older generation, but it must actually check that generation; merely attaching it to a request achieves nothing.
The enforcement belongs where the side effect happens. Checking permission once, then making an unguarded write much later, leaves a race.
Put the ownership check in the write
Here is a small PostgreSQL pattern for a result stored on the job row itself:
UPDATE report_jobs
SET result = $4::jsonb,
status = 'completed'
WHERE id = $1
AND owner_id = $2
AND generation = $3
AND status = 'running'
AND lease_until > clock_timestamp()
RETURNING id;The parameters are the job ID, attempt ID, claimed generation, and result. Assume id is a primary key, lease_until is a non-null timestamptz, and every claim or renewal updates this same row atomically. A new claim increments the generation; renewal preserves it. Workers retain the generation returned by their own claim, rather than reading a newer one and adopting it.
The write returns one row only when its conditions match. No returned row means the worker did not complete the job through this statement. It must not fall back to an unconditional write. It may be stale, expired, or replaying an already completed attempt; inspect durable state to distinguish those cases.
The result and completion state change together. In PostgreSQL's default Read Committed isolation level, an update that waits behind another update rechecks its condition against that updated row. See the transaction isolation documentation.
Keep this transaction short and wait for a successful commit before reporting completion. The deadline predicate is a check during execution, not a promise that commit happens before the deadline. PostgreSQL's date/time documentation explains why clock_timestamp() differs from transaction-start time. This design also assumes the database clock is managed; wall-clock jumps still affect expiry decisions.
This is a guarded completion statement, not a complete queue implementation. It does not protect a separate email, object upload, or external API call. If the lease lives in Redis while the result lives in PostgreSQL, checking Redis and then writing unconditionally to PostgreSQL recreates the same race. The result boundary needs its own enforceable ownership or idempotency rule.
Retrying should converge on one result
Suppose B commits the report, but loses its connection before acknowledging the queue message. The message can arrive again even though the result already exists. A lease alone does not resolve that ambiguity.
Give the logical job a stable identifier and each attempt its own identifier. Use the stable identifier to recognize repeated work; use the attempt identifier to explain ownership and logs. Do not generate a new logical identity on each retry.
For example, a unique constraint on a report's logical key can prevent duplicate result rows. It does not stop an unconditional overwrite or a second email. If publication calls an external API, use the provider's idempotency support where available. A durable outbox can record delivery intent alongside the result, but its sender can still deliver twice if it crashes after sending and before recording success. Deduplication must reach the recipient or provider boundary too.
Recovery also needs a retry policy. Distinguish failures that may resolve from failures that need corrected input. Use bounded backoff with jitter, a retry budget, and a terminal state that someone can inspect. Retrying at every layer can multiply load when a dependency is already struggling; the AWS Builders' Library discussion of retries is useful context.
The goal is not to claim that a job executes exactly once. It is to define which observable effects may repeat, and enforce the ones that must not.
Record the transitions
The most useful records describe a job being claimed, renewed, completed, expired, or rejected as stale. Include the job ID, attempt ID, ownership generation, and reason for the transition. Keep report contents and other sensitive payloads out of those logs.
During an incident, those records answer a practical question: are we waiting for legitimate work, recovering abandoned work, or letting stale attempts change the result?
Test the uncomfortable cases
- Pause A past expiry and let B claim before A resumes. A's completion must change no rows.
- Let A's lease expire without a replacement. A must still fail the deadline check.
- Run two completions for the same current claim. Only one may transition the job to completed.
- Lose the response after a committed result. Redelivery must discover completion without repeating a protected side effect.
- Lose a renewal response. An uncertain worker must not assume permission lasts indefinitely.
- Race renewal, reclaim, and completion. They must obey the same ownership rules.
- Reject a permanent input error. Attempts must stop at the retry policy's limit.
A successful test checks the stored result and external effects, not just whether a worker logged an error.
Design the recovery path as part of the job
Before shipping a background task, ask three questions: when does its claim expire, where can an old attempt be rejected, and what happens if a successful effect is delivered again?
Those answers connect the queue, worker, database, and delivery boundary. Reliability lives in that connection. The job is not finished when it works once; it is finished when the system can recover without trusting a worker that should have lost its authority.