Temporal: Durable Execution for Distributed Systems
Why retry logic, timers, and state machines scattered across your services are a code smell — and how Temporal lets you write distributed workflows as plain, durable functions.
The problem: your business logic is leaking into infrastructure
Any sufficiently complex backend ends up with workflows that span multiple services and need to survive crashes, deploys, and long waits: "charge the card, then wait up to 7 days for the warehouse to confirm shipment, then send the invoice." Implementing this reliably with queues and cron jobs means hand-rolling retries, idempotency keys, and a state machine to track "where is this particular order right now" — usually in a database table that becomes its own maintenance burden.
Temporal's pitch is simple: write that workflow as an ordinary function in your programming language, and Temporal makes it durable — surviving process crashes, deploys, and infrastructure failures — without you writing any of the persistence or retry plumbing yourself.
Workflows as code, not as config
Unlike BPMN-based engines, Temporal workflows are just code — Go, Java, TypeScript, Python, .NET. You write normal control flow (loops, conditionals, function calls) and Temporal's SDK replays the workflow's event history to reconstruct state after any failure.
// Simplified Go workflow
func OrderWorkflow(ctx workflow.Context, order Order) error {
err := workflow.ExecuteActivity(ctx, ChargeCard, order).Get(ctx, nil)
if err != nil { return err }
err = workflow.ExecuteActivity(ctx, ReserveInventory, order).Get(ctx, nil)
if err != nil {
workflow.ExecuteActivity(ctx, RefundCard, order)
return err
}
workflow.Sleep(ctx, 7*24*time.Hour) // durable timer, survives restarts
return workflow.ExecuteActivity(ctx, SendInvoice, order).Get(ctx, nil)
}
That workflow.Sleep call for seven days doesn't hold a goroutine or a server open — Temporal persists the timer in its event history and resumes the workflow exactly where it left off, even if every worker process restarts in the meantime.
Activities: where the side effects live
Workflow code must be deterministic (no direct network calls, no random numbers, no wall-clock reads) because it gets replayed. Actual side effects — API calls, database writes, sending emails — live in activities, which Temporal automatically retries with configurable backoff on failure.
Signals, queries, and long-running state
Workflows can receive external signals (e.g. "customer cancelled order") at any point during execution, and expose queries so other services can inspect a running workflow's current state without disturbing it. This makes Temporal a natural fit for anything stateful and long-lived: subscription billing, saga-style distributed transactions, approval chains, or deployment pipelines that wait on manual gates.
Temporal vs. Camunda vs. a plain queue
- Plain queue / cron — fine for fire-and-forget jobs with no cross-service state to track.
- Camunda — best when business stakeholders need to read/author the process as a visual BPMN diagram.
- Temporal — best when the workflow logic is complex, branchy, and owned entirely by engineers who'd rather express it in code than a diagram tool.
Operational considerations
Temporal Server is a stateful cluster (frontend, history, matching, worker services) backed by a database (Cassandra, PostgreSQL, or MySQL) — plan for it like any other stateful platform component. Temporal Cloud removes the self-hosting burden if you'd rather not run the server fleet yourself. Either way, your own worker processes (running your workflow/activity code) scale horizontally and independently, and deploy like any other service in your stack.
Final thoughts
Temporal's biggest shift is mental, not technical: instead of thinking "how do I persist state and handle retries," you write the workflow as if failures don't exist, and let the runtime handle durability. For engineers who've hand-rolled a saga pattern or a state machine table before, it removes an entire category of infrastructure code.