tencent cloud

Cloud Native Intelligent Gateway

Cache Policy

Download
Focus Mode
Font Size
Last updated: 2026-09-22 18:40:41
AI-Translated

Scenarios

The AI caching capability is built on a two-tier caching architecture (exact cache L1 + semantic cache L2) to help enterprises reduce LLM invocation costs and improve response speed. It is applicable to the following scenarios:
Reduce costs: Requests that hit the cache are not processed by calling the LLM, saving all Token costs. In typical scenarios, this can reduce costs by 30% to 60%.
Improve speed: Exact cache hit latency < 2 ms, and semantic cache hit latency < 200 ms, representing a significant improvement compared to LLM inference (1-10 seconds).
Semantic matching: Requests that are semantically similar but phrased differently can hit the cache, overcoming the limitations of traditional exact matching.
Two-tier architecture: It employs a two-stage query process (exact matching followed by semantic matching) to balance speed and hit rate.

Prerequisites

Precise Cache Prerequisites

1. The AI Gateway instance has been created, the model API has been created, and at least one route has been configured.
2. The Redis service source has been configured.
3. The Redis instance is accessible.

Semantic Cache Prerequisites (Additional)

You must have a Tencent Cloud VectorDB instance and create a database and a Collection within it. The requirements are as follows:
Index type: HNSW
Similarity method: COSINE (Required)
Built-in Embedding: Enabled and a model (such as bge-base-zh) is selected.
Index status: ready

Operation Steps

Configuring Cache

Exact caching uses Redis for storage and returns cached responses for identical LLM requests, skipping LLM calls. Semantic caching uses vector similarity matching and returns cached responses for requests that are semantically similar but phrased differently.

Step 1: Go to the Model API Cache Policy Page

1. Log in to the Microservices Platform console. In the left sidebar, click AI Gateway > Instance List.
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, choose Model Management > Model API.
4. Click the target model API name to go to the details page.
5. Click the Caching Policy Tab. In the Caching Policy Tab, view the current caching configuration status. If it is not configured, Not Configured is displayed.
6. Click Configure Caching Policy to go to the caching configuration dialog box.

Step 2: Configure Precise Cache

After enabling the Whether to Enable switch, configure the following parameters:
Configuration Parameter Description:
Parameter
Required
Description
Recommended Value
Cache key policy
Yes
Cache matching scope
Select "Latest User Message" for single-turn Q&A, and select "Historical Dialogue Mode" for multi-turn conversations.
TTL (seconds)
Yes
Cache validity period, ranging from 60-604800 seconds
3600 (1 hour)
Redis storage configuration
Yes
Cache storage address
Service address, port, username, and password

Step 3: Configure Semantic Cache (Optional)

In the caching configuration dialog, enable the Whether to Enable switch to enable semantic caching.
Configuration Parameter Description:
Parameter
Required
Description
Recommended Value
VDB service source
Yes
Select a pre-configured Tencent Cloud VDB service source
VDB instance in the same region
Similarity Threshold
Yes
The value ranges from 0-1. A similarity score higher than this value is considered a hit.
0.85
Maximum Token distance
Yes
Semantic caching is not used when the total number of request tokens exceeds this value.
50
TTL (seconds)
Yes
Cache validity period, ranging from 60-604800 seconds
3600 (1 hour)
Database
Yes
The system automatically pulls the list of databases under the VDB instance.
-
Vector storage configuration
Yes
Cache storage address
Service address, port, username, and password
Advanced Configuration:
Parameter
Description
Isolate Cache by User
Disabled by default. Independent caches are generated for identical requests from different users to enhance privacy.
Cache Successful Responses Only
Enabled by default. Only successful responses with a 200 status code are cached. Error requests are not cached.
Return cache identifier
Disabled by default. The X-Cache: HIT/MISS identifier is added to the response Header.

Step 4: Save Cache Configuration

Click OK to save the configuration. The cache takes effect immediately.

Viewing Cache Policy List

In the left sidebar, choose Cost Management > AI Caching Configuration to view the caching policy list.

Viewing Cache Hit Statistics

On the Cost Management > AI Cache Hit Statistics page, you can view observable metrics such as cache hit rate and cost savings to evaluate caching effectiveness.

Step 1: Go to the AI Cache Hit Statistics Page

1. Log in to the Microservices Platform console. In the left sidebar, click AI Gateway > Instance List.
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, choose Cost Management > AI Cache Hit Statistics.

Step 2: Configure Filter Conditions

Filtering Condition
Description
Option
Consumer group
Filter by consumer group (all is optional)
All / Specified consumer group
Model API
Filter by model API (all is optional)
All / Specified model API
Time Range
Select the time window for data display.
Today/ This Week/ This Month/ Last 7 Days/ Last 30 Days/ Custom

Step 3: View Core Metric Cards

Quickly View Core Cache Hit Metrics:
Metric Value
Description
Total number of requests
Total number of LLM requests received by the gateway within the selected time range
Exact cache hit count
L1 exact cache hit count, which displays the actual hit count and its percentage.
Semantic cache hit count
L2 semantic cache hit count, which displays the actual hit count and its percentage.
Gateway cache hit rate
(Exact hits + Semantic hits) / Total requests × 100%
Gateway cache cost savings
Estimated cost savings from cache hits that avoid LLM calls (calculated based on model unit price)

Step 4: View Semantic Cache Similarity Distribution

It displays the distribution of similarity scores for semantic cache hits to help evaluate the appropriateness of the semantic cache threshold setting.
Similarity Distribution Analysis:
Similarity Range
Description
0.9 - 1.0
A high similarity hit indicates that the user request is very close to the cached content.
0.8 - 0.9
A medium similarity hit indicates a request with similar semantics but different wording.
0.7 - 0.8
A low similarity hit. Pay attention to whether a mismatch occurs.
0.6 - 0.7
A low similarity hit. It is recommended to check whether a misjudgment occurs.
0 - 0.6
An extremely low similarity hit. There may be a risk of mismatch.
Tuning Reference:
If a large number of hits are concentrated in the interval below 0.8 (accounting for > 30%), it is recommended to appropriately increase the similarity threshold. If the number of hits in the interval above 0.9 is too low (accounting for < 20%), it is recommended to appropriately lower the threshold to improve the hit rate.

Step 5: View Cache Hit Trends and Distribution

Trend and Distribution Panel Description:
Panel Name
Description
Cache hit rate trend
Shows the trend of cache hit rate changes within the selected time range (by hour/day granularity).
Cache type distribution
Displays the hit ratio of exact cache and semantic cache in a pie chart.
Response time comparison
Compares the average response time of cache-hit requests and cache-miss requests, and displays the improvement ratio.

Step 6: View Cache Hit Details

The cache hit details table displays detailed hit data categorized by model API and by consumer.
Column Name
Description
Consumer group
Consumer group to which a request belongs
Cache hit quantity
Number of requests with cache hits (exact + semantic)
Number of cache misses
Number of requests that missed the cache and directly called the LLM
Hit rate
Cache hit count / (hit count + miss count) × 100%
Cost savings
Estimated cost savings from cache hits (¥)
Tokens saved
Estimated number of tokens saved by cache hits
Average response time for cache hits
Average response time of cache-hit requests (ms)
Average response time for cache misses
Average response time of cache-miss requests (ms)
Performance improvement ratio
Cache-miss average response time / cache-hit average response time

Step 7: Export Statistical Reports

On the cache hit statistics page, you can click the Export CSV button to download detailed data under the current filter conditions.

Help and Support

Was this page helpful?

Help us improve! Rate your documentation experience in 5 mins.

Feedback