Workhorse
Getting started

Core concepts

The mental model behind every job — identity, live state, ownership, delivery, and evidence.

Most queue bugs are really model bugs: you assumed a job runs once, or that a dead worker's writes stop mattering, or that finished jobs slow down the queue. This page gives you the model that makes Workhorse's behavior predictable — where a job lives, who may touch it, and what remains afterward.

A job is three kinds of row

Workhorse splits every job across three relations, one per kind of fact.

job stores immutable identity: queue, type, payload, attempt budget, and policy. PostgreSQL writes this row when it accepts the job and never changes its semantic fields.

job_runtime is the live row. It exists only while the job is scheduled, blocked, ready, or active, and it is the only mutable lifecycle relation — workers and SQL functions change it as the job becomes eligible, gains an owner, or returns for another attempt.

job_outcome is the terminal row. When the job becomes succeeded, failed, or canceled, PostgreSQL deletes the runtime row and inserts the outcome in the same transaction.

A committed job therefore has one runtime row or one outcome row, never both. No reader can observe a job as live and terminal at the same time, and terminal jobs leave the indexes workers search — which is why a million finished jobs do not slow the next claim.

The seven states

blocked ──▶ scheduled ──▶ ready ──▶ active ──▶ succeeded
   │                        ▲ ▲        │  └───▶ failed
   └────────────────────────┘ └(retry)─┘  └───▶ canceled
  • blocked: the job waits on prerequisite jobs; its dependency policy decides whether it is released or terminally settled when they finish.
  • scheduled: the job has a future runAt, a retry delay, or a durable sleep in progress.
  • ready: the job is eligible and waiting for a worker slot, in queue-local FIFO order.
  • active: one worker owns the job under a lease and is running its handler.
  • succeeded, failed, canceled: terminal. Result, error, or cancellation envelope is stored.

Admin.getJob returns a snapshot with state, currentAttempt, result, and error, whichever side of the split the job is on.

const admin = new Admin(pool);
const job = await admin.getJob(jobId);
console.log(job?.state, job?.currentAttempt, job?.result);

Who owns active work

When a Worker has free slots, it calls claim_many_v1 with that count. PostgreSQL applies the claim_v1 queue-policy transition to each member, then records the worker ID, a lease expiry, and a fence token on job_runtime. The worker sends every active lease through heartbeat_many_v1 while handlers run.

If heartbeats stop, recovery closes the old attempt and makes the job eligible again. Every semantic write carries the worker ID and fence token, so a stale worker cannot complete, fail, or change a job after a newer worker has claimed it — the full mechanics live in Workers.

Why a handler can run twice

The handler runs outside the transaction that grants ownership. A process can finish an external effect — the email sent, the card charged — and die before PostgreSQL records success. Recovery cannot tell that apart from a crash before the effect, so it must run the handler again.

That is at-least-once delivery, and it shapes how you write handlers:

  • If a stage must not repeat within a job, wrap it in HandlerContext.checkpoint. A completed checkpoint replays its stored result on the next run instead of executing again.
  • If an external call must not repeat across systems, give the provider a stable idempotency key. Checkpoints narrow the duplicate window; only the provider can close it.
worker.handle("invoice.send", async (payload: { invoiceId: string }, ctx) => {
  return ctx.checkpoint("send", () =>
    emailProvider.send({ idempotencyKey: `invoice:${payload.invoiceId}` }),
  );
});

Durable execution covers checkpoints, durable sleeps, and their rules.

Where history goes

Execution leaves append-only evidence, kept off the dispatch path so audits never compete with claims.

job_event records every lifecycle change; attempt_history records how each logical attempt closed. Admin.getJobTimeline merges both into one cursor stream, so you can answer "what happened to job X" from SQL-backed records rather than log lines.

job_query is a separate bounded routing projection for operator reads. Admin.listJobs pages through its immutable creation order, then reads each job's current state from job_runtime or job_outcome. Claims and other lifecycle changes don't rewrite the projection or maintain broad state indexes on job_runtime. Both history relations are partitioned by day and retired by retention policy, independently of live dispatch.

Inside a handler, HandlerContext.checkpoint(name, operation) records a named completed stage. The shorter context.checkpoint form in examples refers to that same durable boundary.

Next

  • Workers — leases, fence tokens, and how a process claims jobs
  • Retries — what happens after an attempt fails
  • Durable execution — preserve completed work across another run

Exact relations, constraints, states, and indexes: architecture reference.