Last updated: October 1, 2026 · Tested with .NET 10 (SDK 10.0.201), built-in Microsoft.AspNetCore.RateLimiting, xUnit + Microsoft.AspNetCore.Mvc.Testing 10.0.12
Short answer: ASP.NET Core has rate limiting built in since .NET 7. Register policies with builder.Services.AddRateLimiter(...), add app.UseRateLimiter(), and attach a policy to an endpoint with .RequireRateLimiting("name") or [EnableRateLimiting("name")]. There are four algorithms: fixed window, sliding window, token bucket and concurrency. I ran all four in a .NET 10 app against a small HTTP client that timestamps every request. The output below is real, and it caught three things most examples get wrong. Rejections are 503 unless you change it. A policy registered with AddFixedWindowLimiter is one counter shared by every client. And the sliding window limiter never sends Retry-After.
Minimal Setup
No NuGet package is needed; the middleware is part of the shared framework.
using System.Threading.RateLimiting;
using Microsoft.AspNetCore.RateLimiting;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddRateLimiter(options =>
{
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
options.AddFixedWindowLimiter("fixed", o =>
{
o.PermitLimit = 5;
o.Window = TimeSpan.FromSeconds(10);
o.QueueLimit = 0;
});
});
var app = builder.Build();
app.UseRateLimiter();
app.MapGet("/fixed", () => "ok").RequireRateLimiting("fixed");
app.Run();
Call UseRateLimiter() after UseRouting() if you call that explicitly (minimal APIs do it for you), so the middleware can read the endpoint's policy. Put it after UseAuthentication() when a policy partitions by user.
Set RejectionStatusCode, or You Get 503
I removed the RejectionStatusCode line and sent six requests:
t= 0.12s /fixed 200
t= 0.12s /fixed 503 Retry-After=10
503 says "the server is failing". It shows up in your error-rate dashboards and 5xx alerts, and clients cannot tell a rate limit from an outage. 429 is the status code defined for this (RFC 6585), and well-behaved clients back off when they see it. Always set it.
Gotcha: A Named Policy Is One Counter for Everybody
This is the one that surprised me. AddFixedWindowLimiter("fixed", ...) (and the sliding, token bucket and concurrency equivalents) does not create a limit per client. It creates one limiter, shared by every caller and every endpoint that uses the policy. I sent five requests from 127.0.0.1, then one from ::1:
127.0.0.1 /fixed 200 (x5)
[::1] /fixed 429
A different client was rejected because the first one used up the shared budget. A second endpoint with the same policy (app.MapGet("/strict", ...).RequireRateLimiting("fixed")) was rejected too, without ever having been called. On a public API, one noisy client using these helpers locks everyone out. Use them for a deliberately global cap, such as protecting a fragile downstream service. For per-client limits, use AddPolicy with a partition key (below).
The Four Algorithms, Measured
Fixed Window: Twice the Limit Across a Boundary
A fixed window counts requests per block of time and resets at the end of each block. With a limit of 5 per 10 seconds, I sent one request at t=0 (this starts the window), four at t=9.6 s and five more at t=10.4 s:
t= 0.00s /fixed 200
t= 9.60s /fixed 200
t= 9.61s /fixed 200
t= 9.61s /fixed 200
t= 9.61s /fixed 200
t= 10.41s /fixed 200
t= 10.41s /fixed 200
t= 10.41s /fixed 200
t= 10.42s /fixed 200
t= 10.42s /fixed 200
Nine requests accepted in 0.8 seconds with a limit of five per ten seconds, all within the rules. Fixed window is fine for small limits (5 login attempts per 15 minutes) where a double burst does no harm. For anything else, use one of the next two.
Sliding Window: Same Traffic, Burst Blocked
options.AddSlidingWindowLimiter("sliding", o =>
{
o.PermitLimit = 5;
o.Window = TimeSpan.FromSeconds(10);
o.SegmentsPerWindow = 5; // 2-second segments
o.QueueLimit = 0;
});
Exactly the same timing:
t= 0.00s /sliding 200
t= 9.62s /sliding 200 (x4)
t= 10.41s /sliding 429 Retry-After=-
t= 10.42s /sliding 429 Retry-After=- (x4 more)
The sliding window still counts the four requests from 9.6 s, so the burst after the boundary is rejected. Note the Retry-After=-: the sliding window limiter does not provide retry-after metadata, so your OnRejected code cannot send that header for it (see the table below).
Token Bucket: Bursts Up to a Cap, Then a Steady Rate
options.AddTokenBucketLimiter("token", o =>
{
o.TokenLimit = 5; // bucket size = max burst
o.TokensPerPeriod = 1; // refill 1 token...
o.ReplenishmentPeriod = TimeSpan.FromSeconds(2); // ...every 2 s
o.QueueLimit = 0;
});
t= 0.00s /token 200
t= 0.07s /token 200
t= 0.08s /token 200
t= 0.08s /token 200
t= 0.08s /token 200
t= 0.08s /token 429 Retry-After=2
t= 0.08s /token 429 Retry-After=2
t= 4.20s /token 200
t= 4.20s /token 200
t= 4.20s /token 429 Retry-After=2
Five requests go through at once (the full bucket), then the client is held to one request per 2 seconds: after 4.2 s exactly two tokens had come back. This matches how real API clients behave, quiet and then bursty, and it is the algorithm I use for most public endpoints.
Concurrency: Limits Requests in Flight, Not per Second
options.AddConcurrencyLimiter("concurrency", o =>
{
o.PermitLimit = 2; // 2 running at once
o.QueueLimit = 2; // 2 more may wait
o.QueueProcessingOrder = QueueProcessingOrder.OldestFirst;
});
app.MapGet("/concurrency", async () => { await Task.Delay(1000); return "ok"; })
.RequireRateLimiting("concurrency");
Six requests at the same moment against an endpoint that takes 1 second:
/concurrency 200 (1107 ms)
/concurrency 200 (1108 ms)
/concurrency 200 (2118 ms) queued, ran second
/concurrency 200 (2118 ms) queued, ran second
/concurrency 429 (121 ms) Reason: "Queue limit reached"
/concurrency 429 (138 ms) Reason: "Queue limit reached"
Two ran, two waited a full second in the queue, two were rejected. Queued requests hold a connection and memory while they wait, so keep QueueLimit small or zero for public traffic. Use this limiter around expensive work (PDF generation, report queries, calls to a slow legacy system), usually in addition to a per-client rate limit.
Which Limiter Sends Retry-After?
| Limiter | Retry-After metadata on rejection | Measured value |
|---|---|---|
| Fixed window | Yes | 10 s (time to window reset) |
| Sliding window | No | none |
| Token bucket | Yes | 2 s (time to the next token) |
| Concurrency | No | none; ReasonPhrase = "Queue limit reached" |
If clients depend on Retry-After, either use fixed window or token bucket, or send a fixed value yourself in OnRejected when the metadata is missing.
Per-Client Limits with AddPolicy
To give every client its own counter, write a policy that returns a partition. Each distinct partition key gets its own limiter. This one partitions by an API key header and puts all anonymous callers into one strict shared partition:
options.AddPolicy("per-key", ctx =>
{
var key = ctx.Request.Headers["X-Api-Key"].ToString();
return string.IsNullOrEmpty(key)
? RateLimitPartition.GetFixedWindowLimiter("anonymous",
_ => new FixedWindowRateLimiterOptions { PermitLimit = 2, Window = TimeSpan.FromSeconds(10) })
: RateLimitPartition.GetFixedWindowLimiter($"key:{key}",
_ => new FixedWindowRateLimiterOptions { PermitLimit = 3, Window = TimeSpan.FromSeconds(10) });
});
/per-key (anonymous) 200
/per-key (anonymous) 200
/per-key (anonymous) 429 Retry-After=10
/per-key alice 200
/per-key alice 200
/per-key alice 200
/per-key alice 429 Retry-After=10
/per-key bob 200 <- unaffected by alice
In production the key is usually the user id from a claim, a tenant id, or the client IP for anonymous traffic. A token bucket per user plus a stricter sliding window per IP for anonymous callers covers most APIs:
options.AddPolicy("per-user", ctx =>
{
if (ctx.User.Identity?.IsAuthenticated == true)
{
var userId = ctx.User.FindFirst("sub")?.Value ?? ctx.User.Identity.Name ?? "unknown";
return RateLimitPartition.GetTokenBucketLimiter($"user:{userId}", _ => new TokenBucketRateLimiterOptions
{
TokenLimit = 200,
TokensPerPeriod = 100,
ReplenishmentPeriod = TimeSpan.FromMinutes(1),
QueueLimit = 0
});
}
var ip = ctx.Connection.RemoteIpAddress?.ToString() ?? "unknown";
return RateLimitPartition.GetSlidingWindowLimiter($"ip:{ip}", _ => new SlidingWindowRateLimiterOptions
{
PermitLimit = 20,
Window = TimeSpan.FromMinutes(1),
SegmentsPerWindow = 4,
QueueLimit = 0
});
});
Keep partition keys bounded (user ids, tenant ids, IPs). Every distinct key creates a limiter in memory, so keying on something like a full URL with ids in it grows without limit.
Behind a Proxy, Every Client Has the Proxy's IP
Behind Nginx, a load balancer, Cloudflare or Azure Front Door, RemoteIpAddress is the proxy, and an IP-partitioned policy turns into one shared bucket again. Run the forwarded-headers middleware first and trust only your own proxy:
builder.Services.Configure<ForwardedHeadersOptions>(o =>
{
o.ForwardedHeaders = ForwardedHeaders.XForwardedFor | ForwardedHeaders.XForwardedProto;
o.KnownProxies.Add(IPAddress.Parse("10.0.0.5"));
});
app.UseForwardedHeaders();
app.UseRateLimiter();
Do not accept X-Forwarded-For from anyone else. A client that can set it gets a new partition, and a fresh limit, on every request.
Global Limiter: Applies on Top of Endpoint Policies
GlobalLimiter runs for every request, including endpoints with no policy. I set a global limit of 3 per 10 seconds and kept the endpoint policy at 5:
options.GlobalLimiter = PartitionedRateLimiter.Create<HttpContext, string>(ctx =>
RateLimitPartition.GetFixedWindowLimiter(
ctx.Connection.RemoteIpAddress?.ToString() ?? "unknown",
_ => new FixedWindowRateLimiterOptions { PermitLimit = 3, Window = TimeSpan.FromSeconds(10) }));
/fixed 200
/fixed 200
/fixed 200
/fixed 429 Retry-After=10 <- global limit hit first
/open 429 Retry-After=10 <- endpoint with no policy, still limited
/health 200 <- .DisableRateLimiting()
Both limits apply, and the stricter one wins. .DisableRateLimiting() also exempts an endpoint from the global limiter. Use it on health and metrics endpoints. A liveness probe that gets a 429 under load makes Kubernetes restart a healthy pod (more in ASP.NET Core health checks on .NET 10).
To combine several global limits (for example per IP and a server-wide concurrency cap), use PartitionedRateLimiter.CreateChained(...). A request must pass all of them.
Controllers
[ApiController]
[Route("api/[controller]")]
[EnableRateLimiting("per-user")]
public class OrdersController : ControllerBase
{
[HttpGet]
public IActionResult List() => Ok();
[HttpPost]
[EnableRateLimiting("token")] // replaces the class-level policy for this action
public IActionResult Create() => Ok();
[HttpGet("export")]
[DisableRateLimiting]
public IActionResult Export() => Ok();
}
Checked with an anonymous client (per-IP limit of 2) and a token bucket of 1:
GET /api/orders 200 200 429 class policy
POST /api/orders 200 429 action policy only, not the exhausted class bucket
GET /api/orders/export 200 200 200 200 200 disabled
A Useful 429 Response
options.OnRejected = async (context, ct) =>
{
var response = context.HttpContext.Response;
if (context.Lease.TryGetMetadata(MetadataName.RetryAfter, out var retryAfter))
response.Headers.RetryAfter = ((int)Math.Ceiling(retryAfter.TotalSeconds)).ToString();
else
response.Headers.RetryAfter = "10"; // sliding window and concurrency give no value
await response.WriteAsJsonAsync(new
{
type = "https://tools.ietf.org/html/rfc6585#section-4",
title = "Too Many Requests",
status = 429
}, options: null, contentType: "application/problem+json", ct);
context.HttpContext.RequestServices.GetRequiredService<ILoggerFactory>()
.CreateLogger("RateLimiting")
.LogWarning("Rate limited {Path} from {IP}", context.HttpContext.Request.Path,
context.HttpContext.Connection.RemoteIpAddress);
};
HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 10
{"type":"https://tools.ietf.org/html/rfc6585#section-4","title":"Too Many Requests","status":429}
Three details from testing this:
OnRejectedruns after the status code has been set fromRejectionStatusCode, so you do not need to set it again.- My first version set
response.ContentType = "application/problem+json"and then calledWriteAsJsonAsync(body, ct). The response went out asapplication/json, becauseWriteAsJsonAsyncsets its own content type. Pass it as thecontentTypeargument, as above. - Use
Math.Ceilingfor the header. Truncating 0.4 seconds toRetry-After: 0invites an immediate retry.
Testing Your Limits with WebApplicationFactory
Rate limits are configuration, and configuration breaks quietly. An integration test catches a missing UseRateLimiter() or a wrong status code. Each WebApplicationFactory starts a fresh app, so limiter state does not leak between tests:
// dotnet add package Microsoft.AspNetCore.Mvc.Testing
// and add "public partial class Program;" at the end of Program.cs
public class RateLimitTests
{
[Fact]
public async Task Sixth_request_in_window_gets_429_with_RetryAfter()
{
await using var factory = new WebApplicationFactory<Program>();
var client = factory.CreateClient();
for (int i = 0; i < 5; i++)
Assert.Equal(HttpStatusCode.OK, (await client.GetAsync("/fixed")).StatusCode);
var rejected = await client.GetAsync("/fixed");
Assert.Equal(HttpStatusCode.TooManyRequests, rejected.StatusCode);
Assert.Equal(TimeSpan.FromSeconds(10), rejected.Headers.RetryAfter?.Delta);
}
[Fact]
public async Task Health_endpoint_is_never_limited()
{
await using var factory = new WebApplicationFactory<Program>();
var client = factory.CreateClient();
for (int i = 0; i < 50; i++)
Assert.Equal(HttpStatusCode.OK, (await client.GetAsync("/health")).StatusCode);
}
}
Passed! - Failed: 0, Passed: 2, Skipped: 0, Total: 2, Duration: 1 s - RateLimitTests.dll (net10.0)
Limits Are per Server Instance
All built-in limiters keep their counters in memory. Three instances behind a load balancer each enforce their own copy, so "100 per minute" becomes up to 300, depending on which instance each request lands on. For a single instance, or where approximate limits are fine, the built-in middleware is enough. For exact limits across instances, enforce them at the gateway (Azure API Management, a YARP gateway with a shared store, Kong, Cloudflare) or use a Redis-backed limiter. The .NET microservices guide shows a YARP gateway in front of several services, which is the natural place for that.
Checklist
RejectionStatusCode = 429.AddPolicywith a partition key for anything per client. TheAdd*Limiterhelpers are one shared counter.- Sliding window or token bucket instead of fixed window, unless the limit is tiny.
QueueLimit = 0on public endpoints.UseForwardedHeaders()withKnownProxiesbefore the limiter when behind a proxy..DisableRateLimiting()on health and metrics endpoints.- A fallback
Retry-Afterfor limiters that do not provide one. - An integration test that proves the 429.
- Limits from configuration, so you can tune them without a redeploy.
FAQ
Do I need a NuGet package for rate limiting in ASP.NET Core? No. Since .NET 7, AddRateLimiter, UseRateLimiter and the four limiters are in the framework (System.Threading.RateLimiting and Microsoft.AspNetCore.RateLimiting).
What status code does the rate limiter return? 503 by default. Set options.RejectionStatusCode = StatusCodes.Status429TooManyRequests.
Is AddFixedWindowLimiter per user or per IP? Neither. It is one counter shared by every client and every endpoint using that policy. Use AddPolicy with RateLimitPartition for per-user or per-IP limits.
Which algorithm should I use? Token bucket for most public APIs, sliding window when you want a strict count over a period, fixed window only for small limits such as login attempts, and concurrency around expensive operations.
Does rate limiting work across multiple servers? Not by itself. The counters are in memory per instance. Use a gateway or a distributed store for global limits.
Conclusion
The middleware is solid and fast, but the defaults and helper methods hide three traps that only show up when you measure: 503 instead of 429, a "per client" limit that is really global, and missing Retry-After on two of the four limiters. Partition by client, pick sliding window or token bucket, exempt infrastructure endpoints, and keep an integration test that proves the 429. If you are hardening an API further, SQL injection prevention in C# and output caching in ASP.NET Core are the next two places I would look.
References: Microsoft: Rate limiting middleware in ASP.NET Core · System.Threading.RateLimiting API reference · RFC 6585: 429 Too Many Requests
This article is part of the ASP.NET Core tutorials guide (APIs, security, observability and deployment).
Comments
Post a Comment