Feature Overview
Weighted routing distributes traffic to multiple model services according to the configured weight ratio. This routing policy is applicable in the following scenarios:
CLB: It distributes traffic to multiple instances of the same model to prevent single-point overload.
A/B Testing: It tests the performance of different vendors or model versions proportionally.
Canary Release: It validates new services with a small amount of traffic and gradually increases the traffic ratio.
Cost Optimization: It directs most traffic to lower-cost services while retaining a small number of high-quality services as a supplement.
Configuration Steps
Step 1: Configuring Weighted Routing in the Model API
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, click Model Management > Model API.
4. Click New or edit an existing API.
5. Select the service type as Multi-Model Service.
6. After completing the Basic Information configuration, go to the Step 2: Select Model Service and Routing Policy page.
7. Select the routing policy as Weighted Routing.
8. Configure a weight for each model service.
Configuration Parameters:
|
Model Service | Select from the list of created model services. | qwen-plus-aliyun
| - |
Weight | The traffic weight for this service, ranging from 1-100 | 70
| Required, integer |
Step 2: Saving and Publishing
After the configuration is complete, click OK to save it. The weighted routing policy takes effect immediately.
Weight Allocation Logic
No limit on the total weight:
The total weight is not required to equal 100. The gateway automatically calculates the proportion.
Service A weight: 50, Service B weight: 30.
Service A traffic proportion = 50 / (50+30) = 62.5%
Service B traffic proportion = 30 / (50+30) = 37.5%
Dynamic Weight Adjustment:
You can edit the model API and adjust the weight configuration at any time. The changes take effect immediately after you save them, without requiring a service restart.
Typical Scenario Examples
Scenario 1: CLB - Multi-Instance Deployment
Requirement: You have deployed three vLLM instances and want to distribute traffic evenly among them.
Configuration:
|
vllm-instance-1 | 33 | 33% |
vllm-instance-2 | 33 | 33% |
vllm-instance-3 | 34 | 34% |
Effect: Each instance receives approximately 1/3 of the traffic, achieving load balancing.
Scenario 2: A/B Testing - New vs. Old Service Comparison
Requirement: Test the newly deployed qwen-max-2.5 model and compare its performance with the existing qwen-max model.
Configuration:
|
qwen-max (legacy version) | 80 | 80% | Primary traffic |
qwen-max-2.5 (new version) | 20 | 20% | Test traffic |
Effect: 20% of the traffic is routed to the new version. After performance data is collected, a decision is made on whether to perform a full cutover.
Scenario 3: Canary Release - Gradual Rollout
Requirement: The new model service requires a canary release, with traffic gradually increasing from 5% to 100%.
Phase 1: Initial Canary (5% Traffic)
|
Legacy service | 95 |
New service | 5 |
Phase 2: Expanded Canary (30% Traffic)
|
Legacy service | 70 |
New service | 30 |
Phase 3: Full Cutover (100% Traffic)
Delete the old service or set its weight to 0.
Scenario 4: Cost Optimization - Multi-Vendor Strategy
Requirement: Most traffic uses cost-effective domestic models, while a small amount of Open AI service is retained as a backup.
Configuration:
|
qwen-max-aliyun | 85 | 85% | Low |
gpt-4o-openai | 15 | 15% | High |
Effect: 85% of the traffic uses the more cost-effective qwen-max, while 15% uses gpt-4o to ensure quality.