Scenarios
Based on the prefix similarity of request prompts, the prefix cache router stably routes inference requests with identical or similar prefixes to the same inference instance, so that the KV Cache already generated on that instance can be continuously hit, thereby skipping the prefill phase and reducing time to first Token (TTFT).
Prerequisites
1. An AI Gateway instance has been created and is in the Running state, with the gateway version ≥ 3.9.5.
2. Multiple model services have been created (at least 2), and the downstream inference service supports KV-Cache (key-value cache).
3. Determine whether the selected model service supports KV-Cache (refer to the inference engine or model vendor documentation for confirmation).
Operation Steps
Step 1: Entering the Model API Configuration Page
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. Choose Model Management > Model API in the left sidebar.
4. Click Create or edit an existing API.
5. After completing the Basic Information configuration, go to Step 2: Select Model Service.
Step 2: Selecting the Service Type and Routing Policy
1. Select Service Type as Multi-Model Service.
2. In the Routing Policy area, select Prefix Cache Routing (Cache-Aware Routing).
Step 3: Configuring the Target Model Service
1. In the Target Model Service area, the system automatically pulls the list of registered inference instances.
2. Select the model services that participate in prefix cache routing by selecting them (multiple selections are allowed, with at least 2 required).
3. The following information is displayed for each service for reference:
Service name (for example, sglang-horizon-01)
Protocol type (custom protocol / openai protocol)
Online status (online/offline)
4. Unselected services are displayed in a semi-transparent state, with a message indicating that they can still receive requests normally but will not perform prefix cache matching logic.
5. The header displays the number of selected items.
|
Model Service Pool | Select from registered model services. Select at least 2. |
Candidate Quantity Range | 2 to 50 |
Step 4: Saving the Configuration
Click OK to save the model API configuration. Prefix cache routing takes effect immediately.
Viewing Runtime Metrics
After the configuration is complete, you can view the runtime metrics on the Routing Policy Tab of the model API details page.
|
Cache hit rate | Hits / total routes, reflecting the KV Cache reuse effect. |
Tokens saved | Number of tokens accumulated and cached on the engine side |
Number of degradations | Number of times degraded to standard routing due to load imbalance |
Must-Knows
1. KV-Cache Capability Dependency: The effectiveness of prefix cache routing depends on whether the downstream inference service supports KV-Cache. For services that do not support it, the configuration can still be saved after selection, but prefix caching will not take effect. Requests will fall back to normal load balancing without any errors.
2. Dynamic Dual-Policy Switching: When a severe load imbalance between instances is detected, the router automatically switches from "prefix matching mode" to "least connections mode" to prevent a single instance from being overloaded. It immediately returns to prefix matching after the load recovers.
3. Oversized Prompt Handling: For prompts exceeding a certain length, the beginning portion is truncated for prefix matching without affecting normal forwarding.
4. Automatic Route Record Maintenance: The gateway automatically maintains the prefix route record table internally, requiring no manual intervention. When an inference instance goes offline or fails a health check, its route records are automatically cleared.
5. Manual Route Record Clearing: The details page supports manually clearing route records for all or specified instances, which is suitable for scenarios such as model hot updates, cache pollution, and traffic switching after canary release. The clearing process does not block requests that are being processed.
6. Integration with Fallback: Prefix cache routing can be used together with global cross-service Fallback. When a candidate service becomes unavailable, Fallback is automatically triggered to switch to a backup service.