tencent cloud

Cloud Native Intelligent Gateway

Rate Limiting Policy

Unduh
Mode fokus
Ukuran font
Terakhir diperbarui: 2026-09-22 18:40:41
Diterjemahkan oleh AI

Scenarios

Rate limiting policies control the frequency of client access to AI Gateway APIs, protecting backend model services from sudden traffic surges. AI Gateway provides a three-tier rate limiting system, which you can flexibly select based on your business requirements:
Rate Limiting Level
Description
Applicable Scenarios
Time Window Policy Supported or Not
Basic Rate Limiting
Global rate limiting, which imposes an overall limit on all requests to the current model API, supports rate limiting based on metrics such as concurrent connections, number of requests, and total number of Tokens.
Protecting the overall capacity of the backend service
Supported
Parameter-based Rate Limiting
Configures rate limiting thresholds based on different dimensions and supports setting rate limiting rules separately for dimensions such as consumer, API, and route.
Multi-tenant scenarios with differentiated quotas for different consumers
Not supported
Fine-grained Rate Limiting
Sets fine-grained rate limiting thresholds for different parameter values (such as client IP address) and matches them in sequence.
Preventing abnormal resource occupation by a single client
Not supported
This document guides you on how to configure and manage rate limiting policies in AI Gateway.

Prerequisites

The AI Gateway instance has been created.
The model API has been created.
A consumer has been created (if you need to apply rate limiting by consumer).

Operation Steps

Rate limiting policies are configured when you create or edit a model API. For details, see Model API.

Step 1: Go to the Target API Rate Limiting Policy Configuration Page

1. Log in to the Microservices Platform console. In the left sidebar, select AI Gateway to go to the instance list.
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, click Model Management, and then click the Model API tab.
4. Click the name of the target model API (the API you have created) to go to its details page, and then click Rate Limiting Policy.

Step 2: Enabling Rate Limiting

Enable the rate limiting switch.

Step 3: Selecting the Rate Limiting Level and Configuring Rate Limiting Rules

In the rate limiting configuration area, first select the rate limiting level. The configuration rules for different levels are as follows.

Basic Rate Limiting

Basic rate limiting applies to the entire API and counts all requests uniformly.
1. Enable basic rate limiting.
2. Select the rate limiting metric type and configure the corresponding threshold based on your actual requirements:
Total Tokens (the total Token consumption allowed within the time window)
Number of Requests (the maximum number of requests allowed within the time window)
Concurrency (the maximum number of concurrent connections allowed)
3. You can click Add Rate Limiting Rule to configure multi-level rate limiting for the current API. For example:
Rule 1: The maximum number of concurrent connections is 1000.
Rule 2: The maximum number of Total Token requests per minute is 1000.
Note:
The threshold for basic rate limiting takes effect as a fallback value when no time window rule is matched. To configure differentiated thresholds for different time periods, see Step 4: Enable the Time Window Policy.

Parameter Rate Limiting

Parameter-based rate limiting allows you to configure rate limiting thresholds separately for different dimensions, with each dimension counted independently.
1. Enable parameter-based rate limiting.
2. Select the rate limiting dimension and configure the corresponding threshold based on your actual requirements:
Consumer: Rate limiting is applied per consumer, with each consumer counted independently. This is suitable for multi-tenant scenarios.
API: Rate limiting is applied per API, with all consumers sharing a quota. This is suitable for protecting backend services.
Route: Rate limiting is applied per route. This applies only to Agent APIs, allowing you to configure different rate limiting policies for different routes.
3. You can click Add Rate Limiting Threshold to configure multi-level rate limiting for consumers. For example:
Rule 1: Consumer A: The maximum number of concurrent connections is 1000.
Rule 2: Consumer A: The maximum number of requests per minute is 5000.

Fine-Grained Rate Limiting

Fine-grained rate limiting allows you to configure rate limiting rules based on specific conditions (such as client IP address) and matches them in the order of configuration.
1. Enable fine-grained rate limiting.
2. Configure rate limiting conditions. Select the condition type and enter the parameter name and value based on your actual requirements.
3. Configure the rate limiting threshold for this condition.
4. You can click Add Rate Limiting Threshold to add multiple rate limiting thresholds for the same condition.
5. You can click Add Rule to add different conditions for fine-grained rate limiting and configure the corresponding rate limiting thresholds.
Attention:
Within the same rate limiting level, you can configure only one rule for the same rate limiting metric. If you attempt to add a duplicate, the system will display a notification and block the operation.
You can simultaneously configure basic rate limiting, parameter-based rate limiting, and fine-grained rate limiting. These three layers of rules take effect in combination. The gateway checks all rules sequentially, and if any rule triggers rate limiting, the request is rejected.

Step 4: Enabling Time Window Policy (Optional)

The Time Window policy allows you to configure differentiated rate limiting thresholds for different time periods, enabling dynamic traffic management during business peak and off-peak hours. Only basic rate limiting supports the Time Window policy. Parameter-based rate limiting and fine-grained rate limiting do not currently support it.
1. In the Basic Rate Limiting configuration area, select the Enable Time Window Policy checkbox.
2. After you select it, the page automatically displays the Time Window configuration area, which includes quick policy options and a list of time window rules.
3. Add time window rules using the following three methods and configure the corresponding rate limiting threshold for the threshold of each rule:
Quick Policy Workday/Non-Workday
Quick Policy Daytime/Nighttime
Custom Time Period
4. After you configure and save the Time Window policy, you can click Rule Effect Preview on the Rate Limiting Policy page to verify the rule matching results. The system then displays the details of the effective basic rate limiting rules for the selected time point, including the effective rule name, priority, rate limiting metric, and threshold.
Note:
When a request arrives, the system determines the effective basic rate limiting threshold according to the following logic:
Scenario
Match Result
Effective Threshold
Current time matches a time window rule.
Match successful
Threshold defined by the matched time window rule
Current time does not match any time window rule.
Match failed
Fallback to the threshold defined in the basic rate limiting configuration (fallback value).
Example: The basic rate limiting configuration allows a maximum of 500 requests per minute. After you enable the Time Window policy, you can add a rule: "Workdays 09:00-18:00: a maximum of 100 requests per minute." Consequently, the limit of 100 requests per minute is applied during workday daytime, and the limit of 500 requests per minute is applied during other time periods.

Step 5: Configuring Custom Response Content (Optional)

By default, when a request triggers rate limiting, the gateway returns a fixed 429 response. If you need to customize the response content to align with the error handling logic on the business side, you can configure it through a rate limiting response policy.
1. At the bottom of the rate limiting policy configuration page, click Advanced Configuration to expand the area. Both Rate Limiting Response Policy and Rate Limiting Response Configuration are located here.
2. Select a rate limiting response policy.
Option
Description
Return directly
Use the gateway's default rate limiting response. No additional configuration is required.
Custom response
Custom Status Codes, Response Body, and Response Headers
3. After you select Custom Response, the page expands the Rate Limiting Response Configuration card. Configure the rate limiting response content.
Default Value
Description
Status code.
HTTP status code returned after rate limiting is triggered, with valid values ranging from 200-599 and a default value of 429.
Request protocol
Displays the request protocol of the current model API (such as OpenAI). This field is read-only and prompts you to write the response body in the error format corresponding to the protocol.
Response Body
Enter the response body content returned after rate limiting is triggered. Click Fill Template with One Click, and the system automatically fills in the standard error response template based on the current request protocol. You can modify it on this basis.
Response header
Click Add Response Header to add a custom response header. If a custom response header has the same name as a default internal header of the gateway, the configuration here takes precedence.
Hide Traffic Throttling Response Headers
Switch. Enabled by default. When the switch is enabled, clients will not receive rate limiting-related response headers.

Step 6: Enabling Request Queuing (Optional)

When a request triggers the rate limiting threshold, the system rejects the request by default and returns a response according to the rate limiting response policy. After you enable request queuing, excess requests enter a waiting queue and are processed when concurrency capacity becomes available, preserving the chance of successful requests when computing resources are constrained.

Enable Request Queuing

At the bottom of the rate limiting policy configuration page, turn on the Request Queuing switch.
Note:
Request queuing takes effect only when a rate limiting rule based on the concurrency dimension is triggered. When rate limiting is triggered based on the request count or total Token count dimension, requests are not queued. This is because exceeding the count limit within the time window means the quota for the current window has been exhausted, and waiting cannot release the counter. In this case, requests are directly returned according to the rate limiting response policy.

Queuing Time Configuration

After you enable Request Queuing, you need to configure the Queue Time. The queue time is divided into three levels based on consumer priority, all measured in seconds. The three levels must satisfy the following condition: high priority ≥ medium priority ≥ low priority.
Default Value
Description
High priority
Maximum queue wait time for high-priority consumers. Default value: 7s. Value range: 0–60s.
Medium priority
Maximum queue wait time for medium-priority consumers. Default value: 3s. Value range: 0–60s.
Low priority
Maximum queue wait time for low-priority consumers. Default value: 1s. Value range: 0–60s.
Note:
1. The queue time indicates the maximum duration that requests at a level are allowed to wait. A longer wait time configured for the high-priority level means that high-value requests are more willing to wait in the queue and less likely to be discarded when computing resources are constrained. A shorter wait time configured for the low-priority level causes low-value requests, such as batch processing requests, to fail fast and release queue positions early. Therefore, the constraint is high priority ≥ medium priority ≥ low priority. Do not set the shortest wait time for the high-priority level.
2. After a request waits longer than the queue time for its level, the system processes it according to the rate limiting response policy configured in Step 5 and returns a 429 response or includes a Retry-After header.

Step 7: Enabling Priority Scheduling (Optional)

After you enable request queuing, the queue is processed by default in first-in, first-out (FIFO) order. After you enable priority scheduling, the queue is dequeued according to the consumer priority level, and high-priority requests are processed first.

Configuring Consumer Priority

Priority scheduling determines the dequeue order based on the consumer priority level. This field is configured in Consumer Management, not on the rate limiting policy page.
1. In the left sidebar, go to Consumer Management.
2. Create a consumer, or click the name of an existing consumer to go to its details page.
3. Configure the Priority field. You can select High, Medium, or Low. Consumers without a configured priority default to Medium. Existing consumers can continue to be used without any changes.

Enable Priority Scheduling

Enable the Priority Scheduling switch.
Note:
Priority scheduling (high-priority requests are dequeued first) takes effect only when queuing is triggered by a global concurrency rule. It does not take effect for consumer-level concurrency rules. Specifically, when a request enters the queue because the global concurrency limit of basic rate limiting is exceeded, the request is dequeued by priority. If a request is queued because the concurrency limit of a specific consumer is exceeded, priority scheduling does not apply.

Step 8: Saving the Configuration

Click OK to save the rate limiting policy configuration.

Bantuan dan Dukungan

Apakah halaman ini membantu?

masukan