# Rate limits for MENA SaaS APIs

A customer imports a catalogue while another waits for an ordinary account screen. Both are within their monthly allowance, yet the shared service is struggling. A requests-per-minute counter cannot explain which operation consumed the scarce resource.

For a SaaS product serving Saudi Arabia and the UAE, rate limits should be an operating contract: who shares a budget, what work consumes it, and what happens when capacity runs out. This guide proposes that contract for a hypothetical product with separately operated Saudi and UAE deployments. Those are deployment boundaries, not claims about required hosting locations or country-specific traffic patterns.

The objective is to protect ordinary requests without letting imports, exports or paid integrations create unbounded work. The decision record and tests below are engineering recommendations, not benchmark results.

*Cover: original SultanByte artwork showing tenant traffic passing through rate, concurrency and spending controls.*

## Start with the resource you need to protect

[OWASP's API4:2023 guidance](https://api-security.owasp.org/editions/2023/en/0xa4-unrestricted-resource-consumption/) treats resource consumption as more than request frequency. It identifies execution time, memory, upload size, batch operations, pagination and third-party spending as limits that can be missing or inappropriate. Its examples include a request that contains many expensive operations despite an endpoint-level rate limit.

Build an operation inventory before choosing an algorithm. For each endpoint, record its expensive work and the earliest point where you can reject it cheaply. A catalogue lookup, a report covering a year of records and an SMS send should not automatically consume identical budgets.

For the hypothetical product, use these separate controls:

*   A request-rate allowance for ordinary API calls, with an explicit short-burst policy.
    
*   A concurrency limit for expensive work already running, such as report generation.
    
*   Bounds on work inside a request: records, files, query complexity and output size.
    
*   A spending reservation for operations that trigger a paid external service.
    

Treat these as independent checks. A request may fit its rate allowance but still be refused because the report workers are occupied. Document whether refused attempts consume any allowance; otherwise customers cannot reconcile the usage they observe.

## Decide whose allowance a request spends

[RFC 6585](https://www.rfc-editor.org/rfc/rfc6585.html#section-4) defines the HTTP 429 response but deliberately leaves user identification and counting policy to the server. The status code does not tell you whether the limit belongs to an IP address, credential, organisation or deployment.

For authenticated business APIs, this proposed design meters the verified tenant and operation class. It can also impose narrower user or integration limits where one actor could consume the organisation's entire allowance. Keep a separate deployment-wide safety limit to protect shared capacity.

Do not let a client-supplied tenant header select a fresh budget without verifying membership. Similarly, issuing another integration credential should not multiply a tenant's purchased allowance unless that is the intended product contract. Preserve the tenant-level budget across credential rotation.

Use IP-based controls as an additional abuse signal rather than the sole measure of a paying customer's activity. In a test fixture, put several legitimate users behind one address and verify that the policy does not confuse that shared address with one business identity.

SultanByte's [tenant-isolation review](https://www.sultanbyte.com/tenant-isolation-for-mena-saas-a-production-review) covers the separate authorisation boundary. Passing a rate-limit check must never grant permission to read an object.

## Choose an algorithm and name its boundary

[Microsoft's throttling guidance](https://learn.microsoft.com/en-us/azure/architecture/patterns/throttling) distinguishes token buckets, leaky buckets, fixed windows and sliding windows. Token buckets permit a configured burst while refilling at a steady rate. Fixed windows are simpler but permit closely spaced bursts on opposite sides of a boundary. Sliding windows address that boundary effect with additional state.

Choose from the workload rather than a popularity ranking. For an interactive API, a token bucket is a reasonable candidate when short bursts are acceptable. For a downstream service that needs smooth arrivals, evaluate paced dispatch. Neither choice removes the need to cap expensive concurrent work.

Write the policy in operational terms. State whether the allowance applies per application instance, deployment or customer across all deployments. Give the counter an owner and specify which requests reach it.

For the Saudi deployment in this example, test its configured allowance using requests spread across all serving replicas. Repeat independently in the UAE deployment. If each replica maintains a full local allowance, scaling out changes the aggregate capacity a caller can consume. That may be acceptable for a documented approximation; it is not the same contract as one shared allowance.

A customer with access to both deployments also needs an explicit answer: two independent allowances, or one coordinated allowance? Do not sell the latter while implementing the former.

## Make counter updates survive interruption

The [Redis INCR documentation](https://redis.io/docs/latest/commands/incr/) provides rate-limiter examples and identifies a failure between incrementing a counter and assigning its expiry. It shows transactions or a Lua script as ways to coordinate the relevant operations in those examples.

The lesson for this design is to review the whole admission operation, not merely confirm that incrementing a number is atomic. Specify how the implementation reads state, checks the allowance, records consumption and sets expiry. Exercise concurrent requests against the actual implementation.

Also define what happens when the counter store is slow or unavailable. For a paid SMS operation, this proposed policy rejects new sends when it cannot reserve capacity safely. For a low-cost read, a bounded local fallback might be acceptable after a documented risk decision. It must have its own ceiling and expiry, with monitoring that shows when it is active.

Avoid a generic exception handler that silently permits everything. Conversely, a universal rejection policy can turn a counter outage into a complete application outage. Choose failure behaviour per operation and test recovery, including whether a restarted or recovered store unexpectedly restores a full allowance.

![Admission controls for a MENA SaaS API: verified tenant context, request allowance, work capacity and spending reservation before execution. Rejections distinguish caller limits from service overload.](https://cdn.hashnode.com/uploads/covers/60ecf4a0fc37a15ec15655e8/916e7bd9-318f-4090-b942-253bd8621e69.png align="center")

*Original infographic: SultanByte. Proposed admission framework informed by OWASP API4:2023, Microsoft throttling guidance and IETF RFCs 6585 and 9110. The diagram describes engineering choices, not measured regional performance.*

## Keep commercial quotas separate from safety controls

[AWS documents API Gateway REST API usage-plan quotas and throttling as best effort](https://docs.aws.amazon.com/apigateway/latest/developerguide/api-gateway-api-usage-plans.html). Clients can exceed configured quotas. AWS explicitly warns against relying on these controls to cap costs or block API access, and says usage-plan API keys should not serve as authentication or authorisation.

That is a product-specific limitation, not a claim that every gateway behaves identically. Before buying or configuring another gateway, establish its documented enforcement scope, approximation and failure behaviour.

For this product, keep commercial usage accounting separate from transient admission counters. Define whether billing counts attempted, accepted or completed work. Use a durable operation identifier so a retry cannot create a second billable operation accidentally.

Where a costly action needs a hard application budget, design reservation before dispatch and reconciliation afterward. Define the unit being reserved and what happens if the provider accepts a request but the response is lost. An uncertain operation should not release its reservation merely because the caller timed out.

This is a proposed accounting workflow, not a promise that an application counter will exactly cap a provider invoice. Delayed charges, other credentials and activity outside this path need their own controls. Use provider-side spending controls where available and alerts to expose discrepancies; an alert alone does not reject a request.

## Return a response clients can act on

Use 429 when the caller has exceeded the applicable request allowance. RFC 6585 says the response should explain the condition, may include `Retry-After`, and must not be stored by a cache.

For temporary service overload, [RFC 9110 defines 503 Service Unavailable](https://www.rfc-editor.org/rfc/rfc9110.html#section-15.6.4). Do not label every counter-store failure as customer overuse. Keep an internal reason code that distinguishes tenant throttling, unavailable admission state and exhausted service capacity.

[RFC 9110's Retry-After definition](https://www.rfc-editor.org/rfc/rfc9110.html#section-10.2.3) permits an HTTP date or a non-negative integer number of seconds. For the proposed API, return a machine-readable error code plus a useful human message. Localise that message for Arabic and English while keeping the code stable. Neither the language nor a translated message should alter retry behaviour.

Only publish a wait estimate that the limiter can justify. When several policies apply, avoid claiming that the request is guaranteed to succeed after one counter resets. Capacity and other applicable limits may still prevent admission.

## Budget retries as part of the workload

[AWS's retry-with-backoff guidance](https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/retry-backoff.html) explains that frequent retries can worsen contention and that retried operations need idempotent behaviour. It also recommends stopping after bounded attempts rather than retrying indefinitely.

For this API, set an overall deadline and attempt budget in the client contract. Honour valid server retry guidance; if the required wait exceeds the remaining deadline, return a pending or failed state rather than retrying early. Use a reviewed backoff policy when the server provides no usable wait.

Assign retry responsibility deliberately. Inspect the SDK, worker and application layers so each does not independently restart the same operation. For a write whose outcome is unknown, reconcile by its durable identifier before issuing a new business operation. A timeout tells the client it lacks an answer, not that the server did nothing.

## A release record your team can use

Keep one policy record per deployment and operation class. Include the principal being metered, algorithm, burst allowance, sustained rate, concurrency ceiling, payload bounds and spending unit. Record the counter location, expiry behaviour, outage policy, response code, retry contract and responsible owner. Mark any approximation plainly.

Then run a small acceptance suite with synthetic tenants:

1.  Send a burst across replicas and around a window boundary. Compare accepted work with the documented policy, including its permitted approximation.
    
2.  Let one tenant saturate report generation while another performs ordinary reads. Check the chosen service objective and tenant-level rejection reasons.
    
3.  Rotate credentials and switch interface language. Confirm neither creates an unintended fresh tenant allowance.
    
4.  Interrupt the counter store during admission. Verify the selected fallback or rejection and its recovery behaviour.
    
5.  Lose a response after a paid operation is accepted. Retry through every enabled layer and inspect execution, reservations and accounting for duplication.
    
6.  Exercise both deployments at once. Confirm the result matches the promised independent or coordinated allowance.
    

Do not approve a higher customer limit from a successful gateway configuration alone. Require evidence that the workers, database and paid dependencies can support it, then retain the policy version with the release. The useful deliverable is a limit the team can explain and enforce through failure, not just a number displayed in an API plan.
