Retry & Dead Letter¶
Handle transient failures automatically. When retries are exhausted, dead-letter handlers provide a fallback.
Global Retry Policy¶
Configure default retries for all jobs.
builder.Services.AddNexJob(options =>
{
options.MaxAttempts = 5; // Initial attempt + 4 retries
});
Default is 10 attempts.
Default delay between attempts¶
Unless a job sets its own delay with [Retry(InitialDelay = ...)], the wait before the next attempt comes from options.RetryDelayFactory. Its default is attempt^4 + 15 seconds plus a random 0-29 s multiplied by (attempt + 1), where attempt is the number of attempts already made (1 after the first failure). That is a polynomial curve with jitter, not an exponential one, and it grows quickly: between 16 and 74 seconds after the first failure and between 4.5 and 7 minutes after the fourth. See Custom Retry Delay to change it.
Per-Job Retry Override¶
Use the [Retry] attribute to override the global policy for specific jobs.
[Retry(5, InitialDelay = "00:00:30", Multiplier = 2.0, MaxDelay = "01:00:00")]
public sealed class ProcessPaymentJob : IJob<PaymentInput>
{
public async Task ExecuteAsync(PaymentInput input, CancellationToken ct)
{
// 5 attempts in total = the first run + 4 retries, with exponential backoff
// (each delay gets ±10% jitter): 30s → 60s → 120s → 240s, never above 1h
}
}
| Parameter | Default | Description |
|---|---|---|
attempts |
Required | Maximum number of attempts (including the first) |
InitialDelay |
None | Delay before the first retry, as a TimeSpan string ("00:00:30"). When it is not set, the global options.RetryDelayFactory computes the delay and Multiplier/MaxDelay are not used |
Multiplier |
2.0 | Exponential backoff multiplier (delay = InitialDelay × Multiplier^(retry − 1)) |
MaxDelay |
No cap | Maximum delay between retries, as a TimeSpan string |
A random ±10% jitter is added to every delay computed from InitialDelay.
Immediate Dead-Letter¶
[Retry(0)] // No retries — go straight to dead-letter on failure
public sealed class WebhookNotificationJob : IJob<WebhookInput>
{
public async Task ExecuteAsync(WebhookInput input, CancellationToken ct)
{
// If this fails, it's immediately sent to dead-letter
}
}
Custom Retry Delay¶
For full control over the timing of jobs that do not set [Retry(InitialDelay = ...)], assign a delegate to options.RetryDelayFactory. It receives the number of attempts already made (1 after the first failure) and returns how long to wait before the next one.
builder.Services.AddNexJob(options =>
{
// Exponential backoff with jitter: 2s, 4s, 8s, ... plus up to 1s
options.RetryDelayFactory = attempt =>
TimeSpan.FromSeconds(Math.Pow(2, attempt))
+ TimeSpan.FromMilliseconds(Random.Shared.Next(0, 1000));
});
Dead-Letter Handlers¶
When all retries are exhausted, the job is marked as Failed and its dead-letter handler is invoked (if registered).
public sealed class PaymentDeadLetterHandler : IDeadLetterHandler<ProcessPaymentJob>
{
private readonly IAlertService _alerts;
private readonly IRefundService _refunds;
public PaymentDeadLetterHandler(IAlertService alerts, IRefundService refunds)
{
_alerts = alerts;
_refunds = refunds;
}
public async Task HandleAsync(
JobRecord failedJob,
Exception lastException,
CancellationToken cancellationToken)
{
// Alert the team
await _alerts.SendAsync(
$"Payment processing failed for job {failedJob.Id} after {failedJob.Attempts} attempts",
cancellationToken);
// Optionally trigger compensation
// var input = JsonSerializer.Deserialize<PaymentInput>(failedJob.InputJson);
// await _refunds.InitiateAsync(input!.OrderId, cancellationToken);
}
}
builder.Services.AddTransient<IDeadLetterHandler<ProcessPaymentJob>, PaymentDeadLetterHandler>();
Safety guarantees:
- Dead-letter handlers run in an isolated DI scope
- Exceptions are logged and swallowed — they never crash the dispatcher
- Handlers are optional — jobs without handlers are simply marked as
Failed
Forwarding a dead-lettered job to Kafka or RabbitMQ¶
Two different things are called "dead-letter" when a broker is involved:
| What failed | Where the dead-letter goes |
|---|---|
| A message that can never become a job (malformed body, missing job type) | The broker's own dead-letter: the Kafka trigger's DeadLetterTopic, a RabbitMQ nack with requeue: false (to the queue's dead-letter exchange, if you configured one), Service Bus dead-lettering, the SQS redrive policy. |
| A job that ran and exhausted its retries, or an Outbox publish that kept failing | Only inside NexJob: the job is Failed in storage and shows in the dashboard, and IDeadLetterHandler<TJob> runs if you registered one. The broker never hears about it. |
The second row is not an oversight. The broker message is acknowledged as soon as it becomes a job, because NexJob storage is the source of truth, and the job runs later, possibly on another node, with retries that can take hours. The broker cannot hold the message that long, so it cannot decide the dead-letter. What you can do is forward a copy to a Kafka topic or a RabbitMQ exchange when the job is dead-lettered, so another service that watches that destination can react (alert, reprocess, audit). The job stays Failed in NexJob.
An open-generic IDeadLetterHandler<> receives the dead-letter of every job type, and the NexJob Outbox publishes the copy, so the forward is durable and retried. This example forwards only the jobs of the trigger's target queue (orders) and sends the original message body, nothing else from the job:
public sealed class OrdersDeadLetterForwarder<TJob>(IScheduler scheduler) : IDeadLetterHandler<TJob>
{
public async Task HandleAsync(JobRecord failedJob, Exception lastException, CancellationToken cancellationToken)
{
// Only jobs of this trigger's target queue, and only plain messages (a trigger stores the body as a string).
if (failedJob.Queue != "orders" || failedJob.InputType != typeof(string).AssemblyQualifiedName)
{
return;
}
// The trigger job stores the message body as a JSON string: recover it verbatim.
var body = JsonSerializer.Deserialize<string>(failedJob.InputJson)!;
await scheduler.EnqueueKafkaAsync(
"orders.dlt",
failedJob.Id.Value.ToString(),
body,
headers: new Dictionary<string, string> { ["x-nexjob-error"] = lastException.Message, }, // optional: leave out to send the body only
cancellationToken: cancellationToken);
}
}
services.AddKafkaProducer(o => o.BootstrapServers = "localhost:9092");
services.AddTransient(typeof(IDeadLetterHandler<>), typeof(OrdersDeadLetterForwarder<>));
For RabbitMQ, publish with EnqueueRabbitMqAsync instead (the message id lets the consumer deduplicate):
await scheduler.EnqueueRabbitMqAsync(
"orders.dlx",
"orders.failed",
body,
messageId: failedJob.Id.Value.ToString(),
cancellationToken: cancellationToken);
Things to know before you use it:
- The queues must be polled.
NexJobOptions.Queuesdefaults todefault. Add the trigger's target queue and the Outbox queue (kafka-producerorrabbitmq-producer), or those jobs are never fetched. - A handler for a specific job type wins. If you also registered
IDeadLetterHandler<OrderJob>, the dispatcher uses that one forOrderJoband the open-generic forwarder is not called for it. The dispatcher calls one handler per job type. - Only the body is kept. The original Kafka key and headers are not stored: NexJob keeps the body, the idempotency key (
kafka:{topic}:{partition}:{offset}, or the RabbitMQMessageId) and the trace parent. - It is not a transaction. The forward is enqueued after the job is marked
Failed. If the process dies in between, that one forward is lost. The Outbox then retries the publish until it succeeds. - The job stays in the dashboard. If you requeue it there and the consuming service also reprocesses the forwarded copy, it is processed twice.
- Do not forward the Outbox publisher itself. Filtering by the trigger's queue keeps it out; forwarding every job would make a failing publish forward itself.
- The body may hold personal data, and forwarding sends it to another system. The error message is optional for the same reason: leave the header out unless the consumer needs it.
This code runs in the repository tests (DeadLetterForwardingTests in NexJob.Kafka.Tests and NexJob.RabbitMQ.Tests). A ready-made forwarder, configured where you add the trigger, is planned.
Failure Data Available¶
The JobRecord passed to dead-letter handlers contains:
Id— the failed job's IDAttempts— how many times it was attemptedMaxAttempts— configured maximumLastErrorMessage— the error message from the last failureLastErrorStackTrace— full stack traceInputJson— the serialized input (deserialize it withSystem.Text.Json)Queue,Tags,CreatedAt,CompletedAt— full execution context
When to Use Retries vs Dead-Letter¶
| Scenario | Approach |
|---|---|
| Transient network error (timeout, connection reset) | Retries with exponential backoff |
| External API rate limiting | Retries with longer delays |
| Data validation error (bad input) | Dead-letter immediately — [Retry(0)] |
| Business rule violation | Dead-letter — retry won't fix it |
| Database deadlock | Retries with short delays (2-3 attempts) |
| Foreign job type (worker lacks job assembly) | Safe deferral (resets attempt, defers without dead-lettering) |
Multi-Service & Foreign Job Safe Deferral¶
When multiple microservices or processes share the same database cluster or storage queue, a worker may dequeue a job whose CLR JobType or InputType belongs to another service and cannot be loaded in the current runtime.
In earlier versions, this resulted in an unhandled type loading exception that consumed retry attempts and eventually poisoned the job into Dead-Letter.
How Safe Deferral Works:¶
- When
DefaultJobInvokerFactorycannot resolveJobTypeorInputType, it raises aForeignJobTypeException. JobExecutorintercepts this exception and treats it as a non-fatal, foreign job event:- Rolls back the attempt increment: Restores
job.Attemptsso the owning service has its full attempt budget. - Defers the job: Reschedules the job at
DateTimeOffset.UtcNow + options.ForeignJobRetryDelay(default 5s) so the worker does not spin in a tight polling loop, giving the owning node an opportunity to pick it up. - Never Dead-Letters:
IDeadLetterDispatcheris never called for foreign jobs.
builder.Services.AddNexJob(options =>
{
// Deferral delay before foreign jobs become visible again for other workers
options.ForeignJobRetryDelay = TimeSpan.FromSeconds(5); // Default: 5s
});
Failed Job Retention & Chunked Purging¶
A job that exhausts its attempts is stored as Failed (its dead-letter handler runs at that moment). Failed jobs stay in storage for troubleshooting and manual re-queuing until JobRetentionService purges them.
Configure retention thresholds and batch sizing in NexJobOptions:
builder.Services.AddNexJob(options =>
{
// How long Failed (dead-lettered) jobs are kept — default 30 days
options.RetentionFailed = TimeSpan.FromDays(30);
// Run the retention purge loop every hour (default)
options.RetentionInterval = TimeSpan.FromHours(1);
// Delete in chunks of 1000 to avoid table locks and WAL / log bloat (default)
options.RetentionBatchSize = 1000;
});
RetentionFailed(default 30 days) decides how longFailedjobs are kept.RetentionDeadLetter(default 60 days) applies toFailedjobs only whenRetentionFailedisTimeSpan.Zero; it also covers rows in the legacyDeadLetterstate written by older versions.- To keep failed jobs indefinitely, set both
RetentionFailedandRetentionDeadLettertoTimeSpan.Zero.
Chunked purging runs across all persistent storage providers (PostgreSQL, SQL Server, Redis, MongoDB).
Anti-Bloat Retention Strategies ([Retention])¶
For ultra high-throughput workloads (for example streaming message triggers) where retaining millions of succeeded jobs until the retention interval causes table and index bloat:
-
Immediate purge (
PurgeOnSuccess = true):Deletes the job as soon as it succeeds. Succeeded jobs will not bloat the[Retention(PurgeOnSuccess = true)] public sealed class FastIngestionJob : IJob<DataChunk> { ... }/jobstables, but lifetime execution statistics are preserved in/catalog. -
Payload stripping (
TrimPayloadOnSuccess = true):Clears[Retention(TrimPayloadOnSuccess = true)] public sealed class LargeDocumentJob : IJob<DocumentPayload> { ... }InputJsonwhen the job succeeds. Metadata (execution duration, completion timestamp, tags, logs) is kept for auditability while cutting the storage footprint.
Note: if a job fails or is dead-lettered, the purge/trim is skipped, so the full diagnostics and input data stay available for retry and debugging.
Next Steps¶
- Throttling — Limit concurrent executions
- Idempotency — Handle retries safely with idempotent jobs
- Common Scenarios — Real-world retry patterns