How Many Virtual Users to Load Test an API
How many virtual users to load test an API: work out the number from peak concurrency, then ramp load until latency or errors break to find the limit.

There is no universal number of virtual users for an API load test, so work it out: multiply your expected peak request rate by the average time each request spends in the system, and the result is the concurrency you need to simulate. Run that number first to confirm the API holds up at expected load. Then keep raising the load in a gradual ramp until latency or error thresholds fail, and treat the last healthy level as your working throughput limit.
The short answer: calculate the number, then ramp until it breaks
A load test can describe its traffic in two ways. Virtual users, often shortened to VUs, simulate concurrent users. Requests per second simulates raw throughput. Those are two views of the same load, and the question of how many virtual users to use is really a question about concurrency: how many requests are in flight at the same moment.
Concurrency is not something you pick from a table. It follows from two numbers you can measure or estimate: how many requests per second the API must serve at peak, and how long each one takes. The next section turns that into arithmetic.
The sequence that works for a launch is short:
- Size the test. Compute the virtual user count from peak request rate and response time.
- Run at expected load. Confirm the API stays healthy at the number you sized for.
- Ramp until it breaks. Increase the load gradually and note the level where latency or errors cross your thresholds. This step is a breakpoint test: load rises in steps until you have found the system’s capacity limit.
The first step answers the question in the title, and the last two answer its second half. Most of the useful information is in the third step, because the sized number only tells you what you were aiming at, not what the API can do.
Calculate the virtual user count from peak concurrency
The formula comes from queueing theory. Little’s law states that the long-term average number of customers in a stationary system equals the long-term average effective arrival rate multiplied by the average time a customer spends in the system. For an API, requests are the customers: concurrency equals request rate multiplied by response time.
A worked example, with made-up numbers: Suppose the API must serve 200 requests per second at peak, and a request currently takes 0.25 seconds on average. Concurrency is 200 multiplied by 0.25, which is 50. A test that holds 50 virtual users, each sending a request as soon as the last one returns, generates that load.
Now change one input. If the API slows to one second per request under pressure, the same 200 requests per second means 200 requests in flight at once, four times as many. That is why the response time you plug in matters as much as the rate: use a measured value from a baseline run or from production monitoring, not a guess, because the answer is wrong by whatever factor the guess was wrong by.
Two more inputs refine the number:
- Pauses between steps. A virtual user modelling a person waits between actions, and a service-to-service caller does not. The arithmetic for adding that pause, and for scripting multi-step journeys, is worked through in the post on realistic multi-step flows.
- Expected peak, not average. Size against the busiest period you expect, since the point of the test is to know the API survives it.
If you would rather skip the conversion, model the load directly as requests per second and let the tool allocate the virtual users it needs. Either way, the number is a starting position that the test is meant to challenge, not a setting to defend.
Why a fixed virtual user count can hide the limit
How your tool schedules virtual users decides whether the test can find a limit at all. The k6 documentation on open and closed models describes two models. In the closed model, VU iterations start only when the last iteration finishes. In the open model, VUs arrive independently of iteration completion.
The difference shows up exactly when it matters. With a fixed pool of virtual users in a closed model, a slowing API holds each user longer, so fewer new requests are sent per second, and the load on the system quietly drops as it degrades. The test then reports a gentler failure than a real crowd would deliver, because real users do not wait politely for the previous one to finish.
An open model keeps sending at the chosen rate whether or not the API keeps up, which is closer to how independent users and callers behave. When the goal is to hold a target throughput, or to push past the point where responses slow down, use arrival-rate scheduling. When the goal is to reproduce a known number of simultaneous sessions, a fixed pool of virtual users is the right model.
This is also why the sizing arithmetic above is only the start. Sized for 50 concurrent requests, a closed test cannot tell you what happens at 200.
Ramp until latency or errors break to find the throughput limit
The throughput limit is the load at which the API stops meeting its targets, and the way to find it is to raise load until something gives. Three rules, which the k6 test-types guide spells out, keep the result readable.
Increase the load gradually. A sudden increase makes it hard to pinpoint why and when the system starts to fail. Keep the shape simple as well: ramp-up, plateau and ramp-down is enough for all test types.
Let thresholds end the run. A breakpoint test commonly has to be stopped, manually or automatically, as thresholds start to fail, and when those problems appear the system has reached its limits. Define the thresholds before the run: a latency percentile you will not exceed and an error rate you will not tolerate. The load at the moment the first one fails is your number.
Keep raising load while the API degrades. Other test types stop once the system degrades past a chosen point; a breakpoint test keeps raising the load through the degradation. A fixed pool of virtual users would taper off as responses slow down, as the previous section explained, so the ramp should raise the arrival rate itself.
Decide what failure means before you start, because it comes in stages: slower responses, then timeouts, then error responses, then collapse. Recording the load at each stage gives you a graded picture and not a single cliff.
One environment warning applies. Avoid breakpoint tests in elastic cloud environments. If autoscaling grows the system as the load rises, the test finds only the limit of your cloud account bill, not of your API. So pin the capacity for the duration of the run or test a fixed-size staging copy.
Repeat the ramp after each fix. Each pass moves the ceiling and shows what the next bottleneck is. Requests built from traffic captured from real usage make the ramp more realistic than a single hand-written call.
Load test, stress test, breakpoint test: what to run before launch
The three names describe different target loads. A load test uses the load you expect on an average day. A stress test pushes load past the expected average to see how the system behaves at its limits. A breakpoint test gradually increases load to identify the capacity limits of the system. A stress test asks whether the API survives a bad day, and a breakpoint test asks where the ceiling is. All three are types of API testing that answer different questions.
For a pre-launch run, use this order:
- A smoke test, the place to start before progressing to higher loads and longer durations.
- An average-load test at the sized virtual user count.
- A stress test above that load. Run stress tests only after average-load tests.
- A breakpoint test, to find the ceiling.
How far above average should the stress step go? There is no fixed percentage, and the right target depends on the stressful situations your API may face. Pick a figure from your own risk, and write it down so the result means something later.
Do not over-read the labels either. A stress test for one application is an average-load test for another, and teams disagree about the names. What counts is that you know which question each run answers.
For teams that would rather not hand-script the traffic, Stresseur describes stress tests that run from one endpoint to ten thousand users. The method above is the same whichever tool generates the load, and k6 and Loadmill compared shows two of them side by side. For a wider comparison, see load testing tools for small teams.
Conclusion
Size the test from peak request rate and response time, run it at expected load, then ramp gradually until thresholds fail. Choose an open, arrival-rate model when you want the load to hold steady while the API slows, so the limit you find is the API’s and not an artefact of the test. Run smoke, average-load, stress and breakpoint tests in that order before launch, and repeat the ramp after every fix.
Frequently asked questions
What is the difference between a load test and a stress test?
A stress test pushes load past the expected average to see how the system behaves at its limits, while a load test targets the load you expect. The two names are relative to the application: a stress test for one application is an average-load test for another.
How do I handle 10,000 concurrent users without performance degradation?
You find out first, by ramping toward that number and watching where thresholds start to fail. A breakpoint test gradually increases load to identify the capacity limits of the system, so it shows whether 10,000 is under or over the ceiling before launch day does.
Is a virtual user the same as a real user?
No. A virtual user is a modelling unit that simulates a concurrent user, and the load can instead be modelled as requests per second to simulate raw throughput. How many real people that corresponds to depends on how long each one pauses between actions.
How much above average load should a stress test go?
There is no fixed percentage. Some testers default to an increase of 50 or 100 percent over average load, but the right target depends on the stressful situations the system may face.