Idempotency Keys: Make API Retries Safe

Build a runnable Python example of safe API retries, then handle concurrent requests, response replay, key expiry, and external side effects.

Idempotency Keys: Make API Retries Safe illustration
On this page8 sections

A client sends an order request, waits, and times out. Did the server reject it, or did it create the order just before the connection broke? A timeout cannot answer that question. Retrying an ordinary create operation can produce a second order.

An idempotency key identifies one intended operation across multiple delivery attempts. The server stores enough information to recognize a retry and return the original result. This guide builds a small executable example, then explains the concurrency, retention, and external-service boundaries that determine whether the design is safe.

Define the contract before writing the handler

The client creates a random key when the user starts an operation, persists it with that operation, and reuses it after network failures. A new purchase gets a new key, even if the basket is identical. Generating a fresh key inside every retry defeats deduplication.

POST /orders
Idempotency-Key: 02da4f91-5797-46bc-9be5-998a3a648b90
Content-Type: application/json

{"sku":"book-42","quantity":2}

Define the server-side scope as an authenticated principal or tenant, an operation name, and the key. Obtain identity from verified authentication, never from a tenant field supplied in the request body. Authentication and authorization must run again before replaying a stored response. A key is a correlation identifier, not a credential.

RequestContract for this example
New scope and keyCreate one order and save the result atomically.
Same key and same validated inputReplay the stored status and response body.
Same key with different inputReject the conflict without performing more work.
Concurrent attempts with the same keySerialize ownership; the later attempt reads the committed result.
Key no longer retainedTreat as a new operation only under the documented retention contract.

Response details are a product decision. For example, Stripe documents replaying the original status and body, including server errors, and checking that reused keys have matching parameters. Do not assume every API caches failures or retains keys for the same period. Document your own behavior alongside the rest of your API design contract.

Keep the result and the write in one transaction

A tempting implementation checks a cache, creates an order, and then saves the response. Two workers can both see a cache miss. A crash after creating the order but before saving the response also leaves the next worker unable to recognize the completed request.

For work confined to one database, put the business write and the idempotency record in the same transaction. Either both commit or neither commits. Return success only after commit. A unique constraint on the key scope protects the storage invariant; the handler must also coordinate concurrent ownership.

Run a local concurrency example

Save this as idempotency_demo.py and run python idempotency_demo.py with Python 3.12 or later and its SQLite module. It creates a disposable database in a temporary directory and runs eight attempts through separate connections. No API keys, packages, or server are needed.

This is the persistence core, not an HTTP server. The caller supplies an already authenticated tenant. Validation admits only a simple SKU and positive integer quantity, then fingerprints that normalized object. More complex APIs must define how defaults, decimals, omitted fields, and API versions affect request identity.

import hashlib
import json
import sqlite3
import tempfile
import uuid
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path


def connect(path):
    # Explicit SQL controls the transaction lifecycle.
    return sqlite3.connect(path, timeout=10, autocommit=True)


def create_order(path, tenant, key, payload):
    if not tenant or not isinstance(key, str) or not 1 <= len(key) <= 128:
        raise ValueError("Invalid tenant or key")
    if not isinstance(payload, dict) or set(payload) != {"sku", "quantity"}:
        raise ValueError("Expected sku and quantity")
    sku, quantity = payload["sku"], payload["quantity"]
    if not isinstance(sku, str) or not 1 <= len(sku) <= 80:
        raise ValueError("Invalid SKU")
    if type(quantity) is not int or not 1 <= quantity <= 100:
        raise ValueError("Invalid quantity")

    normalized = json.dumps(
        {"sku": sku, "quantity": quantity}, sort_keys=True, separators=(",", ":")
    )
    fingerprint = hashlib.sha256(normalized.encode("utf-8")).hexdigest()
    scope = (tenant, "create-order:v1", key)
    db = connect(path)
    try:
        # SQLite serializes writers before the lookup.
        db.execute("BEGIN IMMEDIATE")
        previous = db.execute(
            "SELECT fingerprint, status, body FROM requests "
            "WHERE tenant = ? AND operation = ? AND request_key = ?", scope
        ).fetchone()
        if previous:
            if previous[0] != fingerprint:
                raise ValueError("Key reused with different input")
            result = (previous[1], json.loads(previous[2]))
        else:
            order_id = str(uuid.uuid4())
            db.execute(
                "INSERT INTO orders VALUES (?, ?, ?, ?)",
                (order_id, tenant, sku, quantity),
            )
            body = json.dumps({"order_id": order_id, "quantity": quantity})
            db.execute(
                "INSERT INTO requests VALUES (?, ?, ?, ?, ?, ?)",
                (*scope, fingerprint, 201, body),
            )
            result = (201, json.loads(body))
        db.execute("COMMIT")
        return result
    except BaseException:
        if db.in_transaction:
            db.execute("ROLLBACK")
        raise
    finally:
        db.close()


if __name__ == "__main__":
    with tempfile.TemporaryDirectory() as directory:
        path = Path(directory) / "orders.sqlite"
        db = connect(path)
        try:
            db.executescript("""
                CREATE TABLE orders (
                    id TEXT PRIMARY KEY, tenant TEXT NOT NULL,
                    sku TEXT NOT NULL, quantity INTEGER NOT NULL
                );
                CREATE TABLE requests (
                    tenant TEXT NOT NULL, operation TEXT NOT NULL,
                    request_key TEXT NOT NULL, fingerprint TEXT NOT NULL,
                    status INTEGER NOT NULL, body TEXT NOT NULL,
                    PRIMARY KEY (tenant, operation, request_key)
                );
            """)
        finally:
            db.close()

        payload = {"sku": "book-42", "quantity": 2}
        with ThreadPoolExecutor(max_workers=8) as workers:
            results = list(workers.map(
                lambda _: create_order(path, "tenant-a", "demo-key", payload),
                range(8),
            ))
        assert all(result == results[0] for result in results)
        try:
            create_order(path, "tenant-a", "demo-key", {**payload, "quantity": 3})
        except ValueError as error:
            assert str(error) == "Key reused with different input"
        else:
            raise AssertionError("Expected a conflicting-input rejection")

        db = connect(path)
        try:
            count = db.execute("SELECT COUNT(*) FROM orders").fetchone()[0]
        finally:
            db.close()
        assert count == 1
        print("8 attempts, 1 order, identical responses, conflicting input rejected")

The expected output is 8 attempts, 1 order, identical responses, conflicting input rejected. The assertions verify response equality and the business-row count, not just that a key was cached. The example stores successful results only. Validation failures happen before the transaction, and transactional failures roll back so a later attempt can try again.

SQLite permits one writer at a time, and BEGIN IMMEDIATE starts a write transaction immediately. This makes the lookup and insertion safe in this local example, but serializes unrelated writes too. A busy timeout can still expire; return a retryable failure rather than treating a lock error as permission to bypass deduplication. This is not a performance model for a high-throughput service.

Port the invariant, not the locking shortcut

In PostgreSQL, a plain SELECT followed by INSERT is not enough to claim a missing key. Use a unique key scope and an atomic insertion to claim it inside the transaction, then write the order and stored result before committing. A conflicting attempt waits or receives a bounded retryable outcome according to your timeout policy.

There is a subtle Read Committed case: INSERT ... ON CONFLICT DO NOTHING can decline an insert because of another transaction even when that row was not visible to the insertion's snapshot. Read the committed result in a subsequent statement, and handle a missing row or transaction retry explicitly. Do not assume a same-statement fallback SELECT always finds the winner. PostgreSQL documents this snapshot behavior.

Keep transactions short and enforce lock and statement timeouts. For long-running jobs, a durable operation resource with pending, succeeded, and failed states is often clearer. Concurrent clients can receive an operation URL instead of waiting for the entire job. Worker leases need ownership checks or fencing so an expired worker cannot overwrite a newer result.

A database transaction cannot undo an external payment

Suppose the handler commits an order and calls a payment provider. A process crash between those steps can leave either an unpaid order or an uncertain payment outcome. Putting the network call inside a local database transaction does not make the remote effect transactional.

One option is to write an outbox event with the order and idempotency result in the same transaction. A worker delivers the event later. That delivery can repeat, so derive a stable downstream idempotency key from the operation and use the provider's documented contract. Reconcile uncertain outcomes using the provider's operation identifier or lookup API. If the provider offers neither deduplication nor reconciliation, automated retries cannot promise a single effect.

The event-driven architecture guide explains related consumer and event-flow patterns. The outbox prevents losing the intent to send; it does not by itself prevent duplicate delivery. Avoid advertising an unrestricted exactly-once guarantee across independent systems.

Retention and replay are part of correctness

Choose a retention period that covers your supported retry and redelivery window, including offline clients and delayed workers. After a record is deleted, the same key can create another order. For business operations that must remain unique longer, also enforce a permanent business identifier, such as a merchant order reference.

The demo deliberately has no cleanup job. A production implementation needs expiry metadata, a bounded cleanup policy, and tests for late requests racing with cleanup. Protect stored responses like other tenant data. Cache only fields needed for replay, avoid secrets in keys or logs, and recheck access before returning an old result. If replay includes headers such as Location, persist or reconstruct them consistently too.

Retry with a deadline, bounded attempts, exponential backoff, and jitter. Deduplication limits business effects; it does not make unlimited retries cheap. Track replay rate, payload conflicts, lock wait time, transaction failures, and oldest pending operation. See the observability guide for connecting these signals to request traces.

Test the failure boundaries

  • Send the same key concurrently and verify one business write and identical stored results.
  • Change the payload while keeping the key and verify a conflict with no new write.
  • Use the same key under a different tenant and verify independent operations and no response leakage.
  • Fail after the business INSERT but before COMMIT, then retry and verify the incomplete transaction left no order.
  • Drop the response after COMMIT, then retry and verify the original operation identifier is returned.
  • Test expired keys, lock timeouts, and worker redelivery against the behavior promised to clients.

Start with one endpoint whose database effects can share a transaction. Write its retry contract, run the concurrency example, and add failure injection around commit. Expand to asynchronous or external work only after you can identify which system owns each effect and how an uncertain outcome will be recovered.

References

Share this article

Stuck on implementation?

Get private, 1-on-1 help with system design, performance, scaling, or any technical challenge.

Book a Session

Related Production Resources

Course

Free learning tracks

Turn this guide into a structured production engineering path.

Lab

Interactive engineering labs

Practice the same ideas through scenario-based simulators.

Reference

Production cheatsheets

Keep the operational commands and checks nearby.

Glossary

Key terms

Review the vocabulary behind the architecture.

Discussion

Questions, corrections, or production notes? Add them here so other learners can benefit.

Comments load on demand

To keep this article fast and private by default, the GitHub-powered discussion loads only when you reach this section.

Prefer GitHub? Open the project discussions directly .

Continue Reading

Related practical guides from the same production engineering path.

Backend 8 min read

PostgreSQL EXPLAIN ANALYZE: Read and Fix Slow Query Plans

Read PostgreSQL query plans with a reproducible SQL lab. Diagnose row estimates, loops, buffers, and sort spills before changing indexes.

PostgreSQL SQL
Backend 11 min read

Rate Limiting Algorithms Explained: Token Bucket, Sliding Window, and Leaky Bucket

Implement three production rate limiting algorithms from scratch in Python - token bucket, sliding window log, and leaky bucket. Understand the tradeoffs and pick the right one for your API.

Rate Limiting Python
AI 8 min read

RAG Evaluation: Measure Retrieval and Answer Quality

Build a small Python RAG evaluation harness, separate retrieval failures from answer errors, and test missing evidence and document permissions.

RAG AI
Backend 22 min read

OAuth2 Private Key JWT: Build Client Authentication Without Shared Secrets

Learn how OAuth2 private_key_jwt replaces shared client secrets with signed JWT client assertions, then build and verify the flow end-to-end in Python.

OAuth2 Private Key JWT