Best Practices¶
Production guidelines for NexJob.
Job Design¶
Keep Jobs Small¶
Each job should do one thing. If your job fetches data, transforms it, calls three APIs, and sends two emails, split it into multiple jobs with continuations.
// Bad: does everything
public sealed class ProcessOrderJob : IJob<OrderInput> { ... }
// Good: focused responsibility
public sealed class ValidateOrderJob : IJob<OrderInput> { ... }
public sealed class ChargePaymentJob : IJob<ChargeInput> { ... }
public sealed class SendConfirmationJob : IJob<ConfirmationInput> { ... }
Make Jobs Idempotent¶
Jobs can be retried or re-executed due to failures, orphan recovery, or manual requeue. Design them to handle duplicate execution safely.
public sealed class ChargeCardJob : IJob<ChargeInput>
{
public async Task ExecuteAsync(ChargeInput input, CancellationToken ct)
{
// Check if already charged before charging
var exists = await _payments.ExistsAsync(input.OrderId, ct);
if (exists) return;
await _payments.ChargeAsync(input.OrderId, input.Amount, ct);
}
}
See Idempotency for strategies.
Use Minimal Input¶
Only include what the job needs to execute. Don't pass entire entities — pass IDs and fetch inside the job.
// Bad: passes entire order
public sealed record ProcessOrderInput(Order FullOrder);
// Good: passes only the ID
public sealed record ProcessOrderInput(Guid OrderId);
Retry Configuration¶
Match Retry Policy to Failure Mode¶
| Failure Type | Retries | Delay | Reason |
|---|---|---|---|
| Transient network error | 3-5 | Exponential, 30s-5min | Usually resolves quickly |
| External API rate limit | 5-10 | Exponential, 1min-1hr | May need to wait for window reset |
| Database deadlock | 2-3 | Short, 1-5s | Resolves on next transaction |
| Validation error | 0 | N/A | Retry won't fix bad input |
Set Deadlines for Time-Sensitive Jobs¶
// Promotional email — useless if delayed by 30 minutes
await scheduler.EnqueueAsync<SendPromoEmailJob>(
deadlineAfter: TimeSpan.FromMinutes(10));
Concurrency¶
Size Workers for Your Workload¶
options.Workers = 20;
- Low (5-10): Few concurrent jobs, low resource usage
- Medium (10-30): Typical workloads, good throughput
- High (30-100): Heavy throughput, ensure external services can handle it
Use Throttling for External Services¶
[Throttle("stripe-api", maxConcurrent: 5)]
public sealed class ChargeCardJob : IJob<ChargeInput> { ... }
Size the database connection pool¶
Your database is usually shared. A node needs about Workers + 5 connections, so set Maximum Pool Size (Max Pool Size on SQL Server) to about Workers + 10 and multiply by the number of nodes before scaling out. See Database connections and pool sizing.
Use Queues for Workload Isolation¶
options.Queues = new[] { "default", "emails", "heavy-compute" };
Deploy separate worker instances with different queue configurations:
- Worker A:
["default", "emails"]with 20 workers - Worker B:
["heavy-compute"]with 5 workers
Multi-Service Architecture & Dedicated Ops Host¶
In multi-service ecosystems sharing a database cluster:
1. Dedicated Ops Host: Run a dedicated dashboard container (AddNexJobStandaloneDashboard with DisableWorkers = true) so operational monitoring does not consume worker threads or take lock slots from backend workers.
2. Dashboard Queue Scoping: Scope the UI via options.Queues = ["serviceA-queue"] so engineering teams only view jobs, metrics, and queues relevant to their bounded context.
3. Foreign Job Safe Deferral: While queue separation (options.Queues) is the recommended best practice, if workers encounter foreign job types, NexJob automatically rolls back attempt counts and defers the job via options.ForeignJobRetryDelay rather than failing or dead-lettering it.
Anti-Bloat Retention Strategies (High-Throughput Workloads)¶
In high-throughput environments (e.g., event streaming via Kafka, SQS, or RabbitMQ processing millions of jobs daily), table growth and storage bloat can become problematic even with scheduled chunked purging.
Use the [Retention] attribute on job classes to enforce immediate pruning or payload stripping upon successful completion:
// 1. Ephemeral Jobs: Delete record immediately upon success
[Retention(PurgeOnSuccess = true)]
public sealed class HighFrequencyTelemetryJob : IJob<TelemetryPayload>
{
public async Task ExecuteAsync(TelemetryPayload payload, CancellationToken ct)
{
// Process telemetry...
// On success, the job row is immediately deleted from storage.
// Catalog lifetime statistics (/catalog) are preserved!
}
}
// 2. Payload Stripping: Keep metadata & logs for auditing, strip heavy JSON payloads
[Retention(TrimPayloadOnSuccess = true)]
public sealed class HeavyReportGenerationJob : IJob<LargeReportInput>
{
public async Task ExecuteAsync(LargeReportInput input, CancellationToken ct)
{
// Process heavy input...
// On success, InputJson is stripped (''), saving massive storage space
// while preserving duration, state, queue, and execution logs.
}
}
| Strategy | Attribute Setting | Storage Impact | Auditability |
|---|---|---|---|
| Default | (None) | Job retained until JobRetentionService runs |
Full job record and payload intact |
| Immediate Purge | PurgeOnSuccess = true |
Zero row bloat on success; the row is deleted right after the success is committed | Succeeded jobs vanish from /jobs; lifetime stats preserved in /catalog |
| Payload Stripping | TrimPayloadOnSuccess = true |
Massive space savings (payload set to empty string) | Job record, state, duration, and logs preserved; payload stripped |
[!NOTE] If a job with
PurgeOnSuccess = truefails, it is never purged immediately: it follows normal retry policies and dead-letter retention so operators can diagnose and requeue errors.
Downstream Outage Protection & Circuit Breakers¶
When communicating with external partners (payment gateways, CRM APIs, shipping providers), outages can quickly cause cascading failures: 1. Retries burn rapidly across thousands of jobs. 2. The failing service gets hammered ("metralhadora" effect), worsening their downtime. 3. Storage gets polluted with dead-lettered jobs that were otherwise completely valid.
Protect Queues with Circuit Breakers¶
Group external-facing jobs into dedicated queues (e.g., payments, erp-sync) and enable the queue-level circuit breaker:
builder.Services.AddNexJob(options =>
{
options.ConfigureQueue("payments", queue =>
{
queue.EnableCircuitBreaker(cb =>
{
cb.ConsecutiveFailuresThreshold = 5;
cb.OpenDuration = TimeSpan.FromSeconds(30);
cb.BackoffMultiplier = 2.0;
cb.MaxOpenDuration = TimeSpan.FromMinutes(10);
cb.RecoveryDuration = TimeSpan.FromMinutes(2);
cb.RecoveryConcurrency = 2; // Anti-thundering herd ramp-up
// Protect downstream API: trips on 5xx, timeouts, 429 Too Many Requests (Rate Limits),
// 401 Unauthorized (expired tokens), and network connection drops, while safely ignoring client bugs (400, 403, 404, 422)
cb.BreakOnTransientHttpErrors(includeAuthErrors: true);
cb.BreakOn<TimeoutException>();
});
});
});
Why Ramp-Up Concurrency Matters¶
When an external API comes back online after an outage, NexJob transitions the queue to Recovering mode rather than immediately unleashing full concurrency. With RecoveryConcurrency = 2, only 2 jobs execute concurrently for the duration of RecoveryDuration. This gentle ramp-up protects recovering external services from an immediate thundering herd crash.
Graceful Shutdown¶
When the host stops, NexJob stops fetching new jobs immediately and waits up to NexJobOptions.ShutdownTimeout (default 30s) for running jobs to finish. Jobs still running after that are cancelled through the CancellationToken passed to ExecuteAsync.
A job that is cancelled by shutdown is not treated as a failure: it is requeued right away, its attempt is not consumed, and it is never sent to a dead-letter handler. An OperationCanceledException thrown while the host is not stopping (for example an internal timeout) is still an ordinary failure and follows your retry policy.
HostOptions.ShutdownTimeout must be greater than NexJobOptions.ShutdownTimeout (recommended: add 10 seconds). The .NET host defaults to 30s (5s on older templates), so with the NexJob default the host can kill the process before the drain finishes.
builder.Services.Configure<HostOptions>(o => o.ShutdownTimeout = TimeSpan.FromSeconds(40));
builder.Services.AddNexJob(o => o.ShutdownTimeout = TimeSpan.FromSeconds(30));
Write cancellable jobs: honour the CancellationToken so they can be interrupted and requeued quickly instead of being abandoned to the orphan watcher.
Monitoring¶
Enable OpenTelemetry¶
Collect traces and metrics from day one. See OpenTelemetry.
Set Up Alerts¶
Alert on:
nexjob.jobs.failedincreases — failed executions (it counts every failed attempt, including ones that will still be retried; the dashboard's Failed count shows the jobs that exhausted their attempts)nexjob.jobs.expiredincreases — deadlines too tight or workers insufficientnexjob.job.durationp99 spikes — jobs getting slower
Use the Dashboard¶
Check the dashboard regularly for:
- Queue depth trends
- Failed job error patterns
- Orphaned jobs (indicates worker crashes)
Production Checklist¶
- [ ] Persistent storage (not InMemory)
- [ ] Adequate
MaxAttemptsfor your failure modes - [ ] Dead-letter handlers for critical jobs
- [ ]
[Throttle]on jobs calling external services - [ ]
deadlineAfterfor time-sensitive jobs - [ ] OpenTelemetry configured and exporting
- [ ] Dashboard enabled with authorization
- [ ] Retention policies set (prevent storage growth)
- [ ] Idempotent jobs (see Idempotency)
- [ ] Health check configured (
NexJobHealthCheck) - [ ]
HostOptions.ShutdownTimeoutgreater thanNexJobOptions.ShutdownTimeout
Next Steps¶
- Writing Tests — Test your jobs
- Common Scenarios — Real-world patterns
- Troubleshooting — Debug production issues