tencent cloud

Cloud Native Intelligent Gateway

Model Service

Download
Focus Mode
Font Size
Last updated: 2026-09-22 18:40:41
AI-Translated

Scenarios

You need to add large model services to AI Gateway so that the gateway can proxy requests to the corresponding model providers, enabling unified access, routing, degradation, and key management. AI Gateway supports adding model services from providers such as Hunyuan, Google Gemini, DeepSeek, Qwen, and OpenAI. This document describes how to add, edit, and delete model services for AI Gateway. In addition, AI Gateway provides lifecycle management for model services: it supports three-state management, allows you to manually take services offline to control traffic access, performs health checks by provider protocol category, automatically removes traffic when an exception occurs and automatically restores it after recovery, and records status changes to the operation log and event center.

Operation Steps

Adding a Model Service

1. Log in to the Microservices Platform console. In the left sidebar, click AI Gateway > Instance List.
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, click Model Management. Then, click the Model Service tab. In the service list, click New.
4. In the Create Model Service window, complete the configuration for the first step, Basic Information.
Parameter
Required or Not
Description
Service Name
Yes
Enter the service name. It can contain up to 60 characters, supporting uppercase and lowercase letters in both Chinese and English, digits, and separators ("-", "_"). It cannot start with a digit or separator, and cannot end with a separator.
Service type
Yes
Fixed to "AI Model Service".
Model Vendor
Yes
Select a model vendor:
Standard providers: Hunyuan, OpenAI, Anthropic, Google-Gemini, DeepSeek, Qwen, Moonshot, Zhipu AI, Baidu Qianfan, Tencent Cloud TI-ONE
MaaS providers: TokenHub, AWS Bedrock, Azure OpenAI, Volcano Ark, Google Vertex AI
Custom Provider
Model Protocol
Yes
Depending on the model protocols supported by the model vendor, it supports OpenAI compatibility and Anthropic compatibility.
Service Address
Yes
Confirm the service address of the model service.
If you select a custom vendor, you can manually enter or associate a service source here.
Manual Entry: Evaluate the validity of the service address yourself and ensure it complies with relevant regulations and agreement requirements. The path concatenation feature is provided. When the service address is configured, specify the path prefix separately, and the system will automatically concatenate the complete request address. Two path concatenation modes are supported:
Automatic Concatenation (Default): The model API request path is concatenated after the Base URL. This mode is suitable for most standard OpenAI-compatible protocol services.
Use Fixed Path: The request path is fixed to the Base URL without concatenation. This mode is suitable for scenarios with fixed endpoints, such as Bedrock and Azure.
Associated Service Source: You can select an existing TKE or Polaris service and configure the specific service source, namespace, container service, and request protocol.
Model Key
No
Select a pre-configured API key for this vendor, or click "New Key" to navigate to the key management page and add one. The gateway will use this key to call the corresponding model API.
Key Usage Policy
No
When multiple keys are configured, define how the keys are used. The default is round-robin, which can balance the load among multiple keys.
Periodic Key Rotation
No
After being enabled, the gateway selects and schedules keys only from those whose creation time falls within the period.
Description
No
The description of this service, which facilitates subsequent management.
Attention:
The large model capabilities provided by AI model services are supplied by third parties. AI Gateway does not directly provide these capabilities. You must evaluate the suitability and reliability of the services yourself and ensure that your usage complies with relevant regulations and agreement requirements. Otherwise, you are responsible for any consequences arising from non-compliance.
5. Configure Advanced Configuration (Optional):
In the Advanced Configuration section, configure parameters such as network timeout, SNI, quota, and Tags for the service.
Timeout Configuration Description:
Parameter
Default Value
Scope
Description
Connection timeout
10000 ms
1 ~ 3600000 ms
The timeout duration for establishing a TCP connection to the backend service. It is recommended to set this based on the model inference time. For fast models, the timeout can be appropriately reduced.
Read timeout
60000 ms
1 ~ 3600000 ms
The timeout duration for reading response data from the backend service
Write timeout
60000 ms
1 ~ 3600000 ms
The timeout duration for sending request data to the backend service
Recommended Timeout Configuration Values:
Model Type
Connection Timeout
Read Timeout
Write Timeout
Description
Fast Inference Model (< 5 seconds)
5000 ms
30000 ms
30000 ms
Such as qwen-turbo, gpt-4o-mini
Standard Inference Model (5 ~ 30 seconds)
10000 ms
60000 ms
60000 ms
Default value, applicable to most models.
Slow Inference Model (> 30 seconds)
15000 ms
120000 ms
120000 ms
Such as deep reasoning models, long text generation.
Streaming response
10000 ms
No limit
30000 ms
Read timeout is not limited in streaming mode.
SNI Configuration Description:
Parameter
Required
Description
SNI
No
The Server Name Indication value sent during TLS handshake. Uses the domain name in the service address as SNI when left blank.
Quota Configuration Description:
Parameter
Required
Description
Example
RPM (requests per minute)
No
The maximum number of requests per minute allowed by the model as specified by the vendor. If left blank, quota-aware Fallback cannot be used in the model API.
1000
TPM (tokens per minute)
No
The maximum Token consumption per minute allowed by the model as specified by the vendor. If left blank, quota-aware Fallback cannot be used in the model API.
100000
Number of concurrent requests
No
The maximum number of concurrent requests allowed by the model at the same time as specified by the vendor. If left blank, quota-aware Fallback cannot be used in the model API.
100
Note:
Quota configuration is used for quota-aware fallback scenarios. When a consumer's remaining quota falls below the threshold, the gateway automatically downgrades requests to a fallback service. If the model service is not configured with RPM/TPM, it cannot participate in quota-aware fallback.
Service Tag Description:
Parameter
Required
Description
Example
Tag Key
No
Classification identifier. It is recommended to use keys with business meaning.
region,env,provider
Tag Value
No
The value corresponding to the key
singapore,production,openai
Note:
Service Tags are used to automatically filter services within the model API through the "Tag-based routing" policy. When adding a new model service, you only need to configure Tags. The service is then automatically discovered and used by matching APIs, without requiring any modifications to the API configuration.
6. After completing the basic information, click Next to go to the Select Model Policy step.
Model Selection Method: This configuration determines how the gateway handles the model parameter in client requests.
Specifying a Model
Passthrough Request Model
The gateway will ignore the model parameter in client requests and uniformly use the model you specify in the "Default Model" section below. This mode is suitable for cost control and high availability scenarios, facilitating unified routing and fallback.
Default Model: When the "Model Selection Method" is set to "Specified Model", you must select a specific model name here.
Model Fallback: When it is enabled, the gateway can automatically switch (Fallback) to other available models based on predefined rules if a request to the 'Default Model' fails, ensuring service high availability.
Fallback Rules: After enabling Fallback, you must select or configure the fallback model list and switching rules for when the primary model is unavailable.
The gateway will directly use the model parameter from the client request and forward it to the vendor. This mode is suitable for scenarios that require clients to flexibly control model selection, such as evaluating request latency or counting Token usage. However, when using pass-through requests, the gateway cannot explicitly identify the model name actually requested by the user. Please ensure that the client passes the correct model name.
If the model name in the user request matches the backend vendor's model name or requires no transformation, you can directly pass through the model name from the user request to the backend service.
If you need to replace the model name in the user request with the model name defined by the vendor, you can configure Model Name Mapping information. It supports exact matching and prefix matching (*), and allows you to add multiple mapping rules.
Request Model Name (The model name requested by the client)
Target Model Name (The model name to be rewritten, which is the model name actually used by the backend vendor)
Advanced Configuration (Optional)
Model Parameter Validation: When model parameter validation is enabled, the gateway will validate whether the model parameter in client requests is within the allowed list.
Allowed Model List: Defines the allowlist of model names that clients are allowed to request.
Validation Failure Handling: Defines the handling policy for when model validation fails, supporting "Return 404" or "Fallback to Default Model".
7. After the configuration is completed, click OK to create the model service.
8. After the addition, the newly added service will appear in the service list. Click Service ID/Name to view detailed service information.

Editing a Service

On the Model Service list page, locate the target service and click Edit in its operation column to modify the service configuration information. After making changes, click OK to save.

Deleting a Service

On the Model Service list page, locate the target service and click Delete in its operation column. The system will then perform a dependency check before deletion.
1. The system will display a pop-up window asking you to confirm the deletion and automatically check whether the service is bound to any other resources, such as a model API.
2. Verify results:
If no dependencies exist, the pop-up window will directly display the service ID and name. Click OK to delete it.
If dependencies exist, the pop-up window will display the message "Resource Deletion Dependency Check Result" below the service information, prompt "Unresolved dependencies exist", and list the specific dependency items.
3. If dependencies exist, you must first remove all listed dependencies. After the dependencies are removed, click the Recheck action in the pop-up window. The system will then perform the check again. Once the check passes and the dependency prompt disappears, click OK to finally delete the service. To cancel the deletion, click Cancel.

Model Service Lifecycle Management

Lifecycle State Description

Model services support the following three states. The model service list adds a Status column and an Online/Offline action:
Status
Description
Receive Traffic or Not
Not launched
Manually take offline
No
Running
Running normally and health check passed.
Yes
Exception
Health check failed. Waiting for recovery or manual intervention.
No
State transition trigger conditions:
Transition
Triggering Method
Description
Not launched → Running
Manually bring online
The user clicks "Bring Online" in the console.
Running → Not launched
Manually take offline
The user clicks "Take Offline" in the console, and the system stops receiving traffic.
Running → Exception
Health check failure
Consecutive N probe failures (N is configurable, defaulting to 3)
Exception → Running
Health check recovery
Probe succeeds, and the system recovers automatically (no manual intervention required).
Exception → Running
Manually bring online
The user confirms that the exception has been resolved and manually triggers the transition.

Bringing a Model Service Online/Offline

1. In the left sidebar, click Model Management, and then click the Model Service tab.
2. In the model service list, view the Status column of each service.
3. Perform actions on the target service:
Online: Click Online in the operation column. The service enters the "Running" state and starts receiving traffic. Model services are online by default.
Offline: Click Offline in the operation column, and then click Confirm Offline in the confirmation dialog box. After the service is offline, it stops receiving new model API requests, associated model APIs are no longer routed to this service, and the service configuration is retained, allowing the service to be brought back online.

Configuring a Health Check

1. On the model service details page, locate the Health Check Configuration section and click Edit.
2. Enable the health check switch and configure the general parameters:
Field
Required
Default Value
Description
Check Interval
Yes
30 seconds
Probe interval
Timeout Time
Yes
5 (seconds)
Timeout for a single probe
Failure Threshold
Yes
3 (times)
Mark as abnormal after N consecutive failures
Recovery Threshold
Yes
1 (time)
Mark as recovered after N consecutive successes
Probe Path
Yes
/v1/models
Configurable based on the model protocol.

Help and Support

Was this page helpful?

Help us improve! Rate your documentation experience in 5 mins.

Feedback