Best Practices
Running NexJob reliably in production requires a small set of deliberate decisions at design time: keeping jobs focused and idempotent, choosing retry policies that match your failure modes, isolating workloads by queue, and configuring graceful shutdown so worker restarts never lose work. This page collects those guidelines in one place.
Job Design¶
Keep Jobs Small and Focused¶
Each job should do one thing. If your job fetches data, transforms it, calls three APIs, and sends two emails, split it into multiple jobs connected with continuations.
// Too broad — hard to retry, hard to test
public sealed class ProcessOrderJob : IJob<OrderInput> { ... }
// Focused — each step can retry independently
public sealed class ValidateOrderJob : IJob<OrderInput> { ... }
public sealed class ChargePaymentJob : IJob<ChargeInput> { ... }
public sealed class SendConfirmationJob : IJob<ConfirmationInput> { ... }
Make Jobs Idempotent¶
NexJob can retry or re-execute jobs due to failures, orphan recovery, or manual requeue. Design every job so running it twice with the same input produces the same result.
public sealed class ChargeCardJob : IJob<ChargeInput>
{
public async Task ExecuteAsync(ChargeInput input, CancellationToken ct)
{
// Guard against double-charging
var exists = await _payments.ExistsAsync(input.OrderId, ct);
if (exists) return;
await _payments.ChargeAsync(input.OrderId, input.Amount, ct);
}
}
Use Minimal Input¶
Pass only what the job needs. Avoid embedding full entity objects in job payloads — pass IDs and fetch inside the job instead.
// Stores the entire Order object in the job queue — fragile and bloated
public sealed record ProcessOrderInput(Order FullOrder);
// Stores only the ID — lean and forward-compatible
public sealed record ProcessOrderInput(Guid OrderId);
Retry Configuration¶
Match your retry policy to the kind of failure you expect. A policy suited for a transient network blip is wrong for a validation error that will never self-heal.
| Failure type | Max retries | Delay strategy | Rationale |
|---|---|---|---|
| Transient network error | 3 – 5 | Exponential, 30 s – 5 min | Usually resolves quickly |
| External API rate limit | 5 – 10 | Exponential, 1 min – 1 hr | May need to wait for a window reset |
| Database deadlock | 2 – 3 | Short, 1 – 5 s | Resolves on the next transaction |
| Validation error | 0 | N/A | Retrying bad input never helps |
Set Deadlines for Time-Sensitive Jobs¶
If a job is useless after a certain point — a promotional email, a real-time notification — set a deadlineAfter so NexJob discards it rather than delivering it late.
await scheduler.EnqueueAsync<SendPromoEmailJob>(
deadlineAfter: TimeSpan.FromMinutes(10));
Queue Design and Workload Isolation¶
Separate workloads into dedicated queues and deploy independent worker pools for each tier. This prevents a spike in one workload from starving another.
// Worker A: handles user-facing and email workloads
options.Queues = new[] { "default", "emails" };
options.Workers = 20;
// Worker B: handles slow or memory-intensive compute
options.Queues = new[] { "heavy-compute" };
options.Workers = 5;
Throttle Jobs That Call External Services¶
Apply [Throttle] to any job that calls an external API or shared resource to prevent overwhelming it during high-concurrency bursts.
[Throttle("stripe-api", maxConcurrent: 5)]
public sealed class ChargeCardJob : IJob<ChargeInput> { ... }
Protect Queues from Downstream Outages with Circuit Breakers¶
When an external service goes down, retries can pile up and hammer it further. Enable the queue circuit breaker to pause a queue automatically and ramp traffic back up gradually when the service recovers.
See the Circuit Breaker guide for the states, selective exceptions and how to reset it.
builder.Services.AddNexJob(options =>
{
options.ConfigureQueue("payments", queue =>
{
queue.EnableCircuitBreaker(cb =>
{
cb.ConsecutiveFailuresThreshold = 5;
cb.OpenDuration = TimeSpan.FromSeconds(30);
cb.BackoffMultiplier = 2.0;
cb.MaxOpenDuration = TimeSpan.FromMinutes(10);
cb.RecoveryDuration = TimeSpan.FromMinutes(2);
cb.RecoveryConcurrency = 2; // anti-thundering-herd ramp-up
cb.BreakOnTransientHttpErrors(includeAuthErrors: true);
cb.BreakOn<TimeoutException>();
});
});
});
Tip
RecoveryConcurrency = 2 means only 2 jobs run concurrently during the recovery window. This gentle ramp-up prevents you from crashing a service that is still fragile right after coming back online.
Multi-Service Architectures¶
In environments where multiple services share a database cluster:
- Set
QueuePrefixexplicitly in each service so its default queue ({prefix}.default) survives renames and refactors; the dashboard of a standalone host needs the same value. See the default queue. - Run a dedicated dashboard container with
DisableWorkers = trueso operational monitoring does not consume worker threads or take lock slots from backend workers. - Scope the dashboard to relevant queues with
options.Queues = ["serviceA-queue"]so each team sees only their own jobs. - NexJob handles foreign job types safely by rolling back attempt counts and deferring via
ForeignJobRetryDelayrather than dead-lettering them. See the Multi-Service guide.
Storage Selection and Connection Pool Sizing¶
Use persistent storage (Postgres, SQL Server, MongoDB, or Redis) in production. InMemory storage loses all jobs on restart and is suitable only for tests and local development.
Size the Database Connection Pool¶
Each worker node needs roughly Workers + 5 database connections. Set the connection pool maximum to Workers + 10 and multiply by your node count before you scale out.
"ConnectionStrings": {
"NexJob": "Host=db;Database=nexjob;Username=app;Password=secret;Maximum Pool Size=30"
}
| Workers | Recommended pool max | 3-node cluster total |
|---|---|---|
| 10 | 20 | 60 |
| 20 | 30 | 90 |
| 50 | 60 | 180 |
Warning
Undersizing the connection pool causes timeouts under load. Oversizing it exhausts database server resources. Start with Workers + 10 and adjust based on observed connection wait times.
Control Storage Growth with Retention Policies¶
In high-throughput environments, successful job rows accumulate quickly. Use the [Retention] attribute to control this.
// Delete the row immediately on success — zero table bloat
[Retention(PurgeOnSuccess = true)]
public sealed class HighFrequencyTelemetryJob : IJob<TelemetryPayload> { ... }
// Keep the row for auditing but strip the heavy input JSON
[Retention(TrimPayloadOnSuccess = true)]
public sealed class HeavyReportGenerationJob : IJob<LargeReportInput> { ... }
| Strategy | Effect | Auditability |
|---|---|---|
| Default (none) | Row retained until the retention service runs | Full record and payload intact |
PurgeOnSuccess = true |
Row deleted immediately on success | Succeeded jobs vanish from /jobs; lifetime stats preserved in /catalog |
TrimPayloadOnSuccess = true |
Payload cleared, row kept | State, duration, queue, and logs preserved; payload stripped |
Note
PurgeOnSuccess = true only applies on success. A failing job always follows your normal retry policy and dead-letter retention so operators can diagnose and requeue it.
Graceful Shutdown¶
When the host stops, NexJob immediately stops fetching new jobs and waits up to ShutdownTimeout for running jobs to finish. Jobs still running after the timeout are cancelled via the CancellationToken passed to ExecuteAsync.
A job cancelled by shutdown is not treated as a failure: it is requeued immediately, its attempt counter is not incremented, and it is never sent to a dead-letter handler.
// HostOptions.ShutdownTimeout must exceed NexJobOptions.ShutdownTimeout
builder.Services.Configure<HostOptions>(o =>
o.ShutdownTimeout = TimeSpan.FromSeconds(40));
builder.Services.AddNexJob(o =>
o.ShutdownTimeout = TimeSpan.FromSeconds(30));
Warning
The .NET host defaults to a 30-second shutdown timeout (5 seconds on older project templates). If NexJob's ShutdownTimeout equals the host's, the host can kill the process before the drain finishes. Always set HostOptions.ShutdownTimeout to at least NexJobOptions.ShutdownTimeout + 10 seconds.
Write cancellable jobs — honour the CancellationToken throughout your async calls so jobs can be interrupted quickly and requeued rather than being abandoned to the orphan watcher.
Monitoring and Alerting¶
Enable OpenTelemetry from day one. NexJob emits traces and metrics for every job execution and queue depth change. To be notified when a job fails for good, see the Alerts guide.
Set up alerts on:
nexjob.jobs.failedincreasing — indicates failed attempts (including ones that will still be retried)nexjob.jobs.expiredincreasing — deadlines are too tight or workers are insufficientnexjob.job.durationp99 spiking — jobs are getting slower
Check the dashboard regularly for queue depth trends, failed job error patterns, and orphaned jobs (which indicate worker crashes).
Production Checklist¶
Persistent storage¶
Use Postgres, SQL Server, MongoDB, or Redis — not InMemory.
Retry policies¶
Set MaxAttempts appropriate to each failure mode. Register dead-letter handlers for critical jobs.
Throttling¶
Add [Throttle] to every job that calls an external service or shared resource.
Queue isolation¶
Separate heavy, slow, or external-facing workloads into dedicated queues with independent worker pools.
Circuit breakers¶
Enable EnableCircuitBreaker on queues that communicate with external partners.
Idempotency¶
Design every job to be safe to run more than once with the same input.
Graceful shutdown¶
Set HostOptions.ShutdownTimeout greater than NexJobOptions.ShutdownTimeout and honour CancellationToken in all jobs.
Connection pool sizing¶
Set Maximum Pool Size to Workers + 10 per node.
Retention policies¶
Add [Retention] to high-frequency job classes to control storage growth.
Deadlines¶
Set deadlineAfter on time-sensitive jobs so stale work is discarded rather than delivered late.
Observability¶
Configure OpenTelemetry, enable the dashboard with authorization, and set up alerts on failure and expiry metrics.
Health check¶
Register NexJobHealthCheck and include it in your /healthz endpoint.