Workhorse
Operating

Maintenance and retention

Workers keep PostgreSQL tidy automatically — promotion, recovery, partitions, and evidence retirement — while you set policy and watch for lag.

A durable queue accumulates: delayed jobs waiting to become ready, leases that expired with their workers, history partitions filling up, and evidence that eventually outlives its usefulness. Workhorse handles all of it as bounded background work that workers drive automatically. PostgreSQL coordinates the passes with advisory locks, so scaling from one worker to fifty never multiplies destructive work — one worker wins each due phase, the rest get cheap no-ops. TypeScript, Python, and Go workers all call run_maintenance_v1, which keeps the slow phase order in PostgreSQL. A deployment therefore needs no TypeScript worker to replenish partitions, roll up statistics, retire history, or prune terminal and worker records.

The phases, and how to watch them

run_maintenance_v1 calls the same versioned SQL functions you can call yourself:

  • queue.tick promotes due jobs to ready and recovers expired leases.
  • queue.prepareHistoryPartitions creates future time partitions before data arrives.
  • queue.rollupStatistics folds raw history into minute summaries, then derives hour and day tiers.
  • queue.retainHistory retires history the statistics rollup no longer needs.
  • queue.pruneTerminalStorage removes expired idempotency bindings, then terminal job bundles.
  • queue.pruneWorkerRegistry removes registrations whose processes stopped refreshing them.

Direct calls are for controlled scripts and tests; in production the workers already make them. To observe each pass, give any worker an onMaintenance callback:

const worker = new Worker(queue, {
  onMaintenance(telemetry) {
    metrics.record("workhorse.maintenance", {
      phase: telemetry.phase,
      rows: telemetry.rowsAffected,
      durationMs: telemetry.durationMs,
    });
  },
});

The WorkerMaintenanceTelemetry result names the phase, affected rows, duration, lock outcome, and any error — enough to chart cleanup throughput in whatever monitoring you already run.

Set retention policy

Retention decides how long evidence lives. queue.syncRetentionPolicy stores your application's defaults as a deploy step:

await queue.syncRetentionPolicy({
  jobIdentityRetentionDays: 90,
  terminalOutcomeRetentionDays: 30,
  jobEventRetentionDays: 14,
  attemptHistoryRetentionDays: 14,
  scheduleOccurrenceRetentionDays: 30,
  statisticsRetentionDays: 365,
});

Sync updates only values an operator has not overridden, so a deploy never silently reverts an incident-time decision. During an incident, queue.overrideRetentionPolicy changes selected values without a deploy, and queue.revertRetentionPolicy restores the application defaults for the settings you name. queue.getRetentionPolicy shows the effective policy.

Before any destructive change, preview it:

const impact = await queue.previewRetentionPolicy({ terminalOutcomeRetentionDays: 7 });
console.log(impact.eligible.terminalJobs, impact.capped.terminalJobs);

The preview counts currently eligible rows per category, size-capped with an explicit capped flag. The dashboard settings page runs the same preview before applying a change.

Statistics use a retention ladder inside that policy window. Minute summaries age into hour summaries, and hour summaries age into day summaries, so long windows keep totals and mergeable wait percentiles without paying minute-level storage forever.

PostgreSQL guards the dependency order: it rejects any RetentionPolicyDefinition that could remove a job's identity before its retained evidence, and a retained job_redrive descendant keeps its source identity available for lineage.

Delete terminal data through the retention functions rather than deleting job rows directly. Direct deletion cascades into retained history, so it can remove evidence before the configured windows expire or statistics consume it.

Set maintenance cadence

queue.syncMaintenancePolicy stores one set of scheduling defaults for the database, including an IANA timezone, a local wall-clock time for the daily history retention pass, and the statistics rollup interval, group limit, and recompute window. Setting the rollup interval to zero opts the whole fleet out of statistics. queue.overrideMaintenancePolicy and queue.revertMaintenancePolicy follow the same override-and-revert model as retention, and the dashboard settings page shows each value's source. Workers check eligibility at different moments, but the database owns the shared due state — extra replicas change nothing.

Detect lag before it hurts

Cleanup that silently stalls becomes a full disk weeks later. queue.health budgets exactly this: a stalled statistics rollup, late retention, or a missing history partition each produces a status.reasons entry with a stable code.

const health = await queue.health();
if (health.status.level !== "healthy") {
  for (const reason of health.status.reasons) alert(reason.code, reason.observed);
}

For continuous export rather than polling, Workhorse emits OpenTelemetry metrics for in-process activity automatically once your application configures an SDK. Database-wide gauges — queue depth, ready-work age, expired leases, fleet capacity — come from a dedicated observer:

import { WorkhorseMetricsObserver } from "@stablemates/workhorse";

const observer = new WorkhorseMetricsObserver(pool, {
  onError: (error) => logger.error({ error }, "workhorse metrics collection failed"),
}).start();

// During service shutdown:
observer.stop();

Run exactly one observer per database, beside a long-lived service — every replica sees the same queues, so multiple observers export duplicate gauges. registerQueueMetrics and queue.queueMetricSnapshot are the lower-level pieces if you integrate by hand.

Next


Exact maintenance phases, retention constraints, work bounds, and health fields: architecture reference.