Scenarios
The latency-priority intelligent routing policy automatically routes traffic to model services with lower latency based on their real-time latency status. It is applicable in the following scenarios:
Cross-Region Multi-Path Routing: In scenarios involving multiple IDCs and multi-cloud dedicated lines, the optimal route is selected based on network latency.
End-to-End Response Experience Optimization: When there are significant performance differences among multiple model backends, backends with faster responses are prioritized.
Sufficient Capacity Scenario: For services with sufficient capacity, such as third-party APIs, the goal is to achieve the lowest latency.
Capacity-Constrained Scenario: For self-built clusters or hybrid scenarios, a balance between load and latency is required.
Prerequisites
1. The AI gateway instance has been created and is in a running state.
2. Multiple model services have been created, and there are latency differences among them.
3. Understand the network latency and performance characteristics of each model service.
Operation Steps
Step 1: Go to the Model API Configuration Page
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. Choose Model Management > Model API in the left sidebar.
4. Click New or edit an existing API.
5. After completing the Basic Information configuration, go to Step 2: Select Model Service.
Step 2: Select the Service Type and Routing Policy
1. Select Service Type as Multi-Model Service.
2. In the Routing Policy section, select Latency-Priority Routing.
Step 3: Select the Latency Scenario
In the Latency Scenario section, select an appropriate latency scenario.
Scenario 1: Optimal Network Latency
Scenario 2: Optimal Model Latency
Best Network Latency is applied in cross-region multi-path routing scenarios to prioritize services with the best network link quality. It includes the following scenarios:
Cross-region multi-path routing: In scenarios involving multiple IDCs and multi-cloud dedicated lines, the optimal network path needs to be selected.
Direct Data Center Connection: The backend services have similar performance, and the difference primarily stems from network latency.
Focus on link quality: The business is sensitive to network latency and has no special requirements for model business attributes.
Best Model Latency is applied in end-to-end response experience optimization scenarios, comprehensively considering the processing latency of backend model services. It includes the following scenarios:
Backend performance varies significantly: There are large differences in processing speeds among different model services (for example, from different vendors or with different specifications).
Focus on end-to-end experience: The business is sensitive to overall response time, and user experience needs to be optimized.
Load affects latency: In self-built cluster scenarios, backend load directly impacts response latency.
Select an appropriate routing policy based on service capacity characteristics:
|
Fast policy | Always selects the service with the lowest latency, without considering load balancing. | Services such as third-party APIs with sufficient capacity and stable latency |
Balanced policy | Balances between latency and load to avoid single-point overload. | Self-built clusters or hybrid scenarios with limited capacity |
Step 4: Configure the Target Model Service
1. In the Target Model Service section, click the Select Service drop-down list.
2. Select the model services that need to participate in latency-priority routing (multiple selections allowed).
3. The system automatically monitors the network latency of these services and prioritizes routing to the service with the lowest latency.
Step 5: Save the Configuration
Click OK to save the Model API configuration. The latency-priority routing takes effect immediately.
Must-Knows
1. Network Latency Monitoring Mechanism: The gateway monitors the latency of each service through periodic probing, and actual routing decisions are based on the most recent probing data.
2. Network Latency Data Timeliness: Latency data has a certain update lag and is not suitable for handling second-level network jitter.
3. Capacity Risk of the Fast Policy: When the fast policy is used, all traffic may be directed to the service with the lowest latency. Ensure that this service has sufficient capacity.
4. Latency Trade-off of the Load Balancing Policy: The load balancing policy trades off some latency performance for CLB. It is suitable for scenarios that do not have high latency requirements.
5. Integration with Fallback: Latency-priority routing can be used in conjunction with global cross-service Fallback. It automatically switches to a backup service when the service with the lowest latency becomes unavailable.