Core concepts
The mental model behind every job — identity, live state, ownership, delivery, and evidence.
Most queue bugs are really model bugs: you assumed a job runs once, or that a dead worker's writes stop mattering, or that finished jobs slow down the queue. This page gives you the model that makes Workhorse's behavior predictable — where a job lives, who may touch it, and what remains afterward.
A job is three kinds of row
Workhorse splits every job across three relations, one per kind of fact.
job stores immutable identity: queue, type, payload, attempt budget, and policy. PostgreSQL
writes this row when it accepts the job and never changes its semantic fields.
job_runtime is the live row. It exists only while the job is scheduled, blocked, ready,
or active, and it is the only mutable lifecycle relation — workers and SQL functions change it as the job
becomes eligible, gains an owner, or returns for another attempt.
job_outcome is the terminal row. When the job becomes succeeded, failed, or canceled,
PostgreSQL deletes the runtime row and inserts the outcome in the same transaction.
A committed job therefore has one runtime row or one outcome row, never both. No reader can observe a job as live and terminal at the same time, and terminal jobs leave the indexes workers search — which is why a million finished jobs do not slow the next claim.
The seven states
blocked ──▶ scheduled ──▶ ready ──▶ active ──▶ succeeded
│ ▲ ▲ │ └───▶ failed
└────────────────────────┘ └(retry)─┘ └───▶ canceledblocked: the job waits on prerequisite jobs; its dependency policy decides whether it is released or terminally settled when they finish.scheduled: the job has a futurerunAt, a retry delay, or a durable sleep in progress.ready: the job is eligible and waiting for a worker slot, in queue-local FIFO order.active: one worker owns the job under a lease and is running its handler.succeeded,failed,canceled: terminal. Result, error, or cancellation envelope is stored.
Admin.getJob returns a snapshot with state, currentAttempt, result, and error, whichever
side of the split the job is on.
const admin = new Admin(pool);
const job = await admin.getJob(jobId);
console.log(job?.state, job?.currentAttempt, job?.result);Who owns active work
When a Worker has free slots, it calls claim_many_v1 with that count. PostgreSQL applies the
claim_v1 queue-policy transition to each member, then records the worker ID, a lease expiry, and a fence token on job_runtime. The worker
sends every active lease through heartbeat_many_v1 while handlers run.
If heartbeats stop, recovery closes the old attempt and makes the job eligible again. Every semantic write carries the worker ID and fence token, so a stale worker cannot complete, fail, or change a job after a newer worker has claimed it — the full mechanics live in Workers.
Why a handler can run twice
The handler runs outside the transaction that grants ownership. A process can finish an external effect — the email sent, the card charged — and die before PostgreSQL records success. Recovery cannot tell that apart from a crash before the effect, so it must run the handler again.
That is at-least-once delivery, and it shapes how you write handlers:
- If a stage must not repeat within a job, wrap it in
HandlerContext.checkpoint. A completed checkpoint replays its stored result on the next run instead of executing again. - If an external call must not repeat across systems, give the provider a stable idempotency key. Checkpoints narrow the duplicate window; only the provider can close it.
worker.handle("invoice.send", async (payload: { invoiceId: string }, ctx) => {
return ctx.checkpoint("send", () =>
emailProvider.send({ idempotencyKey: `invoice:${payload.invoiceId}` }),
);
});Durable execution covers checkpoints, durable sleeps, and their rules.
Where history goes
Execution leaves append-only evidence, kept off the dispatch path so audits never compete with claims.
job_event records every lifecycle change; attempt_history records how each logical attempt
closed. Admin.getJobTimeline merges both into one cursor stream, so you can answer "what happened
to job X" from SQL-backed records rather than log lines.
job_query is a separate bounded routing projection for operator reads. Admin.listJobs pages
through its immutable creation order, then reads each job's current state from job_runtime or
job_outcome. Claims and other lifecycle changes don't rewrite the projection or maintain broad
state indexes on job_runtime. Both history relations are partitioned by day and retired by
retention policy, independently of live dispatch.
Inside a handler, HandlerContext.checkpoint(name, operation) records a named completed stage.
The shorter context.checkpoint form in examples refers to that same durable boundary.
Next
- Workers — leases, fence tokens, and how a process claims jobs
- Retries — what happens after an attempt fails
- Durable execution — preserve completed work across another run
Exact relations, constraints, states, and indexes: architecture reference.