Test Multi-Step API Flows With Auth Tokens

Capture an authentication token from a login step, reuse it in every later request in a multi-step API flow, and test that an expired token gets rejected too.

Mustafa BayramogluMustafa BayramogluAbstract illustration of a glowing token passed forward through a chain of connected nodes, with one broken path where a token is rejected

Testing a multi-step API flow that needs authentication means capturing the token a login or auth request returns, storing it as a variable, and sending that value back in the Authorization header of every later request in the chain. The same test carries forward any other value a step returns, like a resource ID a create call hands to a later step in the flow. A complete test also tries the flow with a token that should not work: expired, altered, or replayed, and checks that the API rejects it. That is a different problem from running the same flow at load with thousands of virtual users, which is a scale question, not a chaining one.

Capture the token, then pass it forward

A login or authentication request in a chained API test returns a token, either as a field in the JSON response body, commonly named access_token, or as a JSON Web Token, or in a response header. The test needs to read that value out of the response and store it as a variable scoped to the run instead of hardcoding it, because a fresh run gets a fresh token every time the flow executes.

Every request that follows in the chain then sends that stored value back, typically as a bearer token in its Authorization header. One worked multi-step API testing example points the header value directly at the prior step’s response field, instead of retyping the token into a fixed string, so the test still works after the token itself changes.

Concretely: a POST /auth/login call returns { "access_token": "abc123" }. The next call in the chain sets Authorization: Bearer abc123 by referencing that stored value, not by pasting the string in. If the flow has a third step, that step reads the same stored variable again; nothing about the token changes between step two and step three unless the API issues a new one.

The part that makes a multi-step test actually work is this store-then-reference pattern, repeated for every protected call in the flow, so the chain keeps working when the token’s value changes from run to run.

Carry IDs forward too, and test the token that should fail

A token is not the only value a chained flow passes forward. A create step commonly returns an ID for the resource it just made, and a later step needs that same ID to confirm the resource exists or to remove it again. The mechanic is the same as the token: read the value out of the response, store it, and reference the stored variable in the next request instead of guessing or hardcoding an ID.

The flow is not fully tested until it is run with a token that should not work. Deliberately inject an expired token, a token with an altered scope, or a token replayed from an earlier, already-completed flow, and confirm the API answers with a 401 Unauthorized or a 403 Forbidden instead of letting the chain continue as if the token were valid. This is one of the API testing basics: a test suite that only ever runs the happy path never exercises the code that is supposed to reject exactly this.

Both checks belong in the same test file as the working flow, not in a separate suite that might get skipped: one flow that succeeds end to end with a valid token and valid IDs, and one or more flows that deliberately break a link in the chain to confirm the API notices.

What sets this up by hand, and what learns it from traffic

Postman collections, Checkly’s multi-step checks, and Datadog Synthetics all give you a built-in mechanism for this: a way to grab a JSON field or a header value out of one response and carry it into the next. A CLI tool takes a different shape but does the same job: Hurl captures a token or any other value out of one response and carries it into a later request in the same file, pulling it out with an XPath query, a JSONPath query, or a regex match against the body. A hand-written suite in Python (pytest plus httpx) or Node (supertest) reads and reuses those values the same way, just in code instead of a Hurl file. Every one of these still asks the person writing the test to name which field to extract and which header to fill in, step by step, for every chain in the suite.

Stresseur starts from a different place: it learns how an API is actually used from real recorded traffic and creates the tests every change needs from that, so nobody writes the flow, login step included, by hand.

This is a correctness test, not a load test

Everything above is about one flow working end to end: the right token in the right header, the right ID handed to the right step, and the wrong token correctly rejected. That is a correctness question, and it stays the same whether the test runs once or a thousand times.

Running that same flow with thousands of concurrent virtual users is a different problem, covered in load testing realistic multi-step flows: at that point the token is no longer a single value to reference correctly but a resource to provision at scale. One shared login collapses at that volume: ten thousand virtual users signing in as the same username collide on a database constraint, or start getting served a cached response that no longer reflects the real system, and that post’s fix is a dedicated account per peak virtual user instead of one login shared by all of them.

Build the correctness test first. A chained flow that does not work correctly for one user will not produce a meaningful result at load either; it will just fail faster, at scale, with a less obvious cause.

What to do next

Store the token a login step returns, reference it in every later request’s Authorization header, and carry forward any other value, such as a resource ID, the same way. Run the flow again with a token that should fail, and check that the API actually rejects it. Whether that setup lives in a point-and-click tool, a CLI config, or a hand-written test suite, the mechanic is the same; the difference is how much of it you write by hand versus how much a tool that learns from real traffic builds for you.

Frequently asked questions

Does the token come from the response body or a response header?

Either, depending on the API: a login response commonly returns the token as a field in a JSON body, such as access_token, or sets it in a response header instead. The test reads whichever the login step actually returns and stores that value as a variable for later requests.

What HTTP status code should an expired or altered token get instead of 200?

A 401 Unauthorized or a 403 Forbidden, not the 200 a valid chain gets. A test that only ever sends a valid token never confirms the API actually enforces that rejection.

Is a chained-flow test the same thing as a load test with many virtual users?

No. A chained-flow test checks that one flow works correctly end to end with the right token and IDs handed between steps; a load test runs that same flow with many concurrent virtual users to find where the API's capacity runs out, which turns the token into an account-provisioning problem instead of a single value to reference.

All posts