Hypothesis & Fuzzing

Modeling REST APIs as State Machines

You have a CRUD REST service and a pile of example tests that each exercise one endpoint, yet production still throws 404s on resources that were just created, or returns a deleted record from a list endpoint. The defect lives in the interaction between requests — create, update, delete, re-read — not in any single handler. Stateful and model-based testing generates those request sequences for you and shrinks any failure to a minimal reproduction, the same way ordinary property-based testing with Hypothesis shrinks a single failing input. This guide maps REST verbs onto a RuleBasedStateMachine: endpoints become rules, created resources flow through Bundles, and business rules become invariants.

REST resource lifecycle as a Hypothesis state machine POST creates a resource whose server-issued id is pushed into the items Bundle. GET and PUT are non-destructive self-loops that draw an existing id from the Bundle without changing which ids are live. DELETE consumes the id, moving the resource to a terminal removed state so it is never drawn again. After every transition an invariant compares GET /items against the in-memory model dict. REST verbs as state-machine transitions POST /items @rule(target=items) push server-issued id Resource live id held in items Bundle drawable by later rules Removed id consumed — never drawn again DELETE consumes(items) GET read PUT update @invariant runs after every transition assert GET /items == in-memory model dict (live ids only)
Each REST verb is a transition: POST pushes a server-issued id into the items Bundle, GET and PUT draw that id non-destructively, and DELETE consumes it to a terminal removed state. The @invariant reconciles the list endpoint against the in-memory model after every step.

The mapping between REST verbs, Hypothesis primitives, and the in-memory oracle is the whole trick:

REST verbHypothesis primitiveEffect on the items Bundle
POST /items@rule(target=items)pushes the server-issued id
GET /items/{id}@rule(item_id=items)draws an existing id (non-destructive)
PUT /items/{id}@rule(item_id=items)draws an existing id (non-destructive)
DELETE /items/{id}@rule(item_id=consumes(items))removes the id permanently
GET /items@invariant()compares the collection to the model dict

Prerequisites

  • Python 3.9+, hypothesis >= 6.0, pytest >= 7.0
  • A web app exposing CRUD endpoints. The example uses FastAPI with starlette.testclient.TestClient (pip install "fastapi>=0.110" httpx), but the pattern applies to any client.
  • Familiarity with the primitives in Stateful and Model-Based Testing@rule, @initialize, Bundle, @invariant, @precondition.

Solution

Once that table is internalized the translation is mechanical: each verb is a transition, the set of live resource ids is a Bundle, and the server's contract is a set of invariants checked against an in-memory model after every step.

Python
# test_items_api_state.py
from fastapi import FastAPI, HTTPException
from fastapi.testclient import TestClient
from hypothesis import strategies as st, settings, HealthCheck
from hypothesis.stateful import (
    RuleBasedStateMachine, Bundle, rule, initialize, invariant,
    precondition, consumes,
)

# --- A minimal system under test -------------------------------------------
app = FastAPI()
_DB: dict[int, dict] = {}
_NEXT = {"id": 1}

@app.post("/items")
def create_item(body: dict):
    item_id = _NEXT["id"]; _NEXT["id"] += 1
    _DB[item_id] = {"id": item_id, "name": body["name"]}
    return _DB[item_id]

@app.get("/items/{item_id}")
def read_item(item_id: int):
    if item_id not in _DB:
        raise HTTPException(status_code=404)
    return _DB[item_id]

@app.put("/items/{item_id}")
def update_item(item_id: int, body: dict):
    if item_id not in _DB:
        raise HTTPException(status_code=404)
    _DB[item_id]["name"] = body["name"]
    return _DB[item_id]

@app.delete("/items/{item_id}")
def delete_item(item_id: int):
    if item_id not in _DB:
        raise HTTPException(status_code=404)
    del _DB[item_id]
    return {"deleted": item_id}

@app.get("/items")
def list_items():
    return list(_DB.values())

# --- The state machine -----------------------------------------------------
names = st.text(min_size=1, max_size=12)

class ItemsAPIMachine(RuleBasedStateMachine):
    items = Bundle("items")  # server-issued ids of live resources

    @initialize()
    def setup(self):
        _DB.clear(); _NEXT["id"] = 1     # reset server state per sequence
        self.client = TestClient(app)
        self.model: dict[int, str] = {}  # id -> expected name

    @rule(target=items, name=names)
    def create(self, name: str):
        resp = self.client.post("/items", json={"name": name})
        assert resp.status_code == 200
        item_id = resp.json()["id"]      # use the REAL server id
        self.model[item_id] = name
        return item_id                   # push id into the `items` Bundle

    @rule(item_id=items, name=names)
    def update(self, item_id: int, name: str):
        resp = self.client.put(f"/items/{item_id}", json={"name": name})
        assert resp.status_code == 200
        self.model[item_id] = name

    @rule(item_id=items)
    def read(self, item_id: int):
        resp = self.client.get(f"/items/{item_id}")
        assert resp.status_code == 200
        assert resp.json()["name"] == self.model[item_id]

    @rule(item_id=consumes(items))       # consumes => id never drawn again
    def delete(self, item_id: int):
        resp = self.client.delete(f"/items/{item_id}")
        assert resp.status_code == 200
        del self.model[item_id]
        # Contract: a deleted resource must now 404.
        assert self.client.get(f"/items/{item_id}").status_code == 404

    @precondition(lambda self: True)
    @invariant()
    def list_count_matches_model(self):
        # Business rule: the collection endpoint lists exactly the live items.
        listed = self.client.get("/items").json()
        assert len(listed) == len(self.model)
        assert {i["id"] for i in listed} == set(self.model)

TestItemsAPI = ItemsAPIMachine.TestCase
TestItemsAPI.settings = settings(
    max_examples=100,
    stateful_step_count=40,
    suppress_health_check=[HealthCheck.too_slow],
)

The resource lifecycle is what the rules are drawn from: each transition below becomes a rule, and each forbidden transition becomes an assertion.

A resource lifecycle expressed as transitions A timeline of a REST resource lifecycle: creation returns an identifier, reads and updates are valid while the resource exists, deletion makes it absent, and any later read must return 404 rather than stale content. A resource lifecycle expressed as transitions POST created id returned, 201 GET / PUT live reads reflect last write DELETE absent 204, then gone GET again 404 forever no resurrection, no stale body
Every arrow is a rule and every impossible arrow is an invariant — that mapping is the whole design of an API state machine.

Why this works

The Bundle is what makes generated request sequences realistic: a read or delete only ever fires against an id the server actually issued from a prior create, so Hypothesis never wastes steps probing random 404s and instead explores the meaningful state space. Using consumes(items) on delete removes the id from the queue, so the deleted-then-read contract is enforced by construction rather than by chance. The in-memory model dict is the oracle — cheap, obviously-correct code that the real handlers must agree with after every transition, which the @invariant checks.

Edge cases and failure modes

  • Shared server state leaks across sequences. Module-level _DB must be reset in @initialize; otherwise resources from a previous example pollute the next and invariants fail spuriously. Prefer a fresh app/database fixture per machine instance for real services.
  • consumes vs plain Bundle reference. Using the bare Bundle (item_id=items) in delete would leave the deleted id drawable, and a later read would assert 200 against a 404 — a false failure. Always consumes() on terminal operations.
  • Non-deterministic ids. If the server issues UUIDs or relies on wall-clock ordering, capture the id from the response body (as above) rather than predicting it; never hardcode item_id=1.
  • Slow real I/O. Hitting a live database for 40 steps across 100 examples is thousands of requests. Suppress HealthCheck.too_slow, lower stateful_step_count, or use an in-process TestClient as shown. See reducing Hypothesis test execution time.
  • Auth and pagination. Endpoints requiring tokens or returning paged lists need the invariant to follow pagination; comparing only the first page against a growing model will fail past the page size. When the endpoint reaches out to a downstream service, stub it with mocked network and HTTP calls so the state machine stays deterministic across sequences.

Keeping API state machines fast and deterministic

An API state machine that talks to a real server is the slowest test in the suite, so two decisions dominate whether it stays in CI: what it talks to, and how it cleans up.

Run against an in-process application instance rather than a network server whenever the framework allows it. httpx.ASGITransport, Django's test client and Flask's test client all dispatch directly into the application, removing the socket, the TLS handshake and the port allocation — typically a tenfold reduction in per-request cost, with no change to the code under test.

State cleanup has to be part of the machine, not of a fixture. A RuleBasedStateMachine runs many sequences per test, and a fixture only tears down once, so leftover resources from sequence one become phantom results in sequence two. Implement teardown() on the machine itself to delete everything the run created, and keep a set of created identifiers so the cleanup is exact rather than a truncate-everything sweep that would break parallel runs.

Python
from hypothesis.stateful import RuleBasedStateMachine, rule

class ApiMachine(RuleBasedStateMachine):
    def __init__(self):
        super().__init__()
        self.client = make_in_process_client()
        self.created: set[str] = set()        # exact cleanup list for this sequence

    @rule()
    def create(self):
        resp = self.client.post("/orders", json={"total": 1})
        assert resp.status_code == 201
        self.created.add(resp.json()["id"])

    def teardown(self):
        for order_id in self.created:          # runs after EVERY sequence
            self.client.delete(f"/orders/{order_id}")

Determinism needs the same treatment as in any other test: freeze the clock, seed the randomness, and make identifiers come from the server rather than from the strategy so two sequences cannot collide. Where the API is genuinely non-deterministic — a rate limiter, an eventual-consistency read — model it explicitly as a set of acceptable outcomes rather than asserting a single one, or exclude that endpoint from the machine and test it directly.

Mapping HTTP semantics to machine rules A table mapping four HTTP behaviours - creation, idempotent update, deletion and conditional requests - to the rule or invariant that expresses each in a state machine. Mapping HTTP semantics to machine rules Criterion Expressed as Asserts POST creates a rule id is new and readable PUT is idempotent a rule run twice second call changes nothing DELETE then GET a rule + invariant 404, never a stale body ETag / If-Match a precondition rule stale write is rejected
Idempotence and conditional requests are the two properties hand-written API tests almost never cover, and both are one rule each here.

Frequently Asked Questions

How do I represent a POST that creates a resource in a Hypothesis state machine? Model the POST as an @rule whose target is a Bundle, returning the created resource id from the response. Hypothesis pushes that id into the Bundle so later GET, PUT, and DELETE rules draw a real, server-issued id rather than fabricating one.

How do I model DELETE so the same id is not reused after removal? Consume the id from the Bundle with consumes(resources) in the DELETE rule. consumes removes the value from the queue, so no subsequent rule will draw a deleted id, matching the server's contract that a deleted resource returns 404.

Where do business rules go in a REST state machine? Encode them as @invariant methods that query the live API and compare against the in-memory model — for example asserting the list endpoint returns exactly the set of created-and-not-deleted ids after every step. Should the state machine run against a real database? Yes, if the API's correctness depends on database constraints — uniqueness, cascade deletes, transaction isolation — because those are exactly the behaviours a model cannot fake. Use a per-worker database and let teardown() remove only what the sequence created. Where the database is only storage and the logic lives in the application, an in-memory repository is faster and finds the same bugs.

← Back to Stateful and Model-Based Testing