Scenarios
You need to add large model services to AI Gateway so that the gateway can proxy requests to the corresponding model providers, enabling unified access, routing, degradation, and key management. AI Gateway supports adding model services from providers such as Hunyuan, Google Gemini, DeepSeek, Qwen, and OpenAI. This document describes how to add, edit, and delete model services for AI Gateway. In addition, AI Gateway provides lifecycle management for model services: it supports three-state management, allows you to manually take services offline to control traffic access, performs health checks by provider protocol category, automatically removes traffic when an exception occurs and automatically restores it after recovery, and records status changes to the operation log and event center.
Operation Steps
Adding a Model Service
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, click Model Management. Then, click the Model Service tab. In the service list, click New.
4. In the Create Model Service window, complete the configuration for the first step, Basic Information.
|
Service Name | Yes | Enter the service name. It can contain up to 60 characters, supporting uppercase and lowercase letters in both Chinese and English, digits, and separators ("-", "_"). It cannot start with a digit or separator, and cannot end with a separator. |
Service type | Yes | Fixed to "AI Model Service". |
Model Vendor | Yes | Select a model vendor: Standard providers: Hunyuan, OpenAI, Anthropic, Google-Gemini, DeepSeek, Qwen, Moonshot, Zhipu AI, Baidu Qianfan, Tencent Cloud TI-ONE MaaS providers: TokenHub, AWS Bedrock, Azure OpenAI, Volcano Ark, Google Vertex AI Custom Provider |
Model Protocol | Yes | Depending on the model protocols supported by the model vendor, it supports OpenAI compatibility and Anthropic compatibility. |
Service Address | Yes | Confirm the service address of the model service. If you select a custom vendor, you can manually enter or associate a service source here. Manual Entry: Evaluate the validity of the service address yourself and ensure it complies with relevant regulations and agreement requirements. The path concatenation feature is provided. When the service address is configured, specify the path prefix separately, and the system will automatically concatenate the complete request address. Two path concatenation modes are supported: Automatic Concatenation (Default): The model API request path is concatenated after the Base URL. This mode is suitable for most standard OpenAI-compatible protocol services. Use Fixed Path: The request path is fixed to the Base URL without concatenation. This mode is suitable for scenarios with fixed endpoints, such as Bedrock and Azure. Associated Service Source: You can select an existing TKE or Polaris service and configure the specific service source, namespace, container service, and request protocol. |
Model Key | No | Select a pre-configured API key for this vendor, or click "New Key" to navigate to the key management page and add one. The gateway will use this key to call the corresponding model API. |
Key Usage Policy | No | When multiple keys are configured, define how the keys are used. The default is round-robin, which can balance the load among multiple keys. |
Periodic Key Rotation | No | After being enabled, the gateway selects and schedules keys only from those whose creation time falls within the period. |
Description | No | The description of this service, which facilitates subsequent management. |
Attention:
The large model capabilities provided by AI model services are supplied by third parties. AI Gateway does not directly provide these capabilities. You must evaluate the suitability and reliability of the services yourself and ensure that your usage complies with relevant regulations and agreement requirements. Otherwise, you are responsible for any consequences arising from non-compliance.
5. Configure Advanced Configuration (Optional):
In the Advanced Configuration section, configure parameters such as network timeout, SNI, quota, and Tags for the service.
Timeout Configuration Description:
|
Connection timeout | 10000 ms | 1 ~ 3600000 ms | The timeout duration for establishing a TCP connection to the backend service. It is recommended to set this based on the model inference time. For fast models, the timeout can be appropriately reduced. |
Read timeout | 60000 ms | 1 ~ 3600000 ms | The timeout duration for reading response data from the backend service |
Write timeout | 60000 ms | 1 ~ 3600000 ms | The timeout duration for sending request data to the backend service |
Recommended Timeout Configuration Values:
|
Fast Inference Model (< 5 seconds) | 5000 ms | 30000 ms | 30000 ms | Such as qwen-turbo, gpt-4o-mini |
Standard Inference Model (5 ~ 30 seconds) | 10000 ms | 60000 ms | 60000 ms | Default value, applicable to most models. |
Slow Inference Model (> 30 seconds) | 15000 ms | 120000 ms | 120000 ms | Such as deep reasoning models, long text generation. |
Streaming response | 10000 ms | No limit | 30000 ms | Read timeout is not limited in streaming mode. |
SNI Configuration Description:
|
SNI | No | The Server Name Indication value sent during TLS handshake. Uses the domain name in the service address as SNI when left blank. |
Quota Configuration Description:
|
RPM (requests per minute) | No | The maximum number of requests per minute allowed by the model as specified by the vendor. If left blank, quota-aware Fallback cannot be used in the model API. | 1000 |
TPM (tokens per minute) | No | The maximum Token consumption per minute allowed by the model as specified by the vendor. If left blank, quota-aware Fallback cannot be used in the model API. | 100000 |
Number of concurrent requests | No | The maximum number of concurrent requests allowed by the model at the same time as specified by the vendor. If left blank, quota-aware Fallback cannot be used in the model API. | 100 |
Note:
Quota configuration is used for quota-aware fallback scenarios. When a consumer's remaining quota falls below the threshold, the gateway automatically downgrades requests to a fallback service. If the model service is not configured with RPM/TPM, it cannot participate in quota-aware fallback.
Service Tag Description:
|
Tag Key | No | Classification identifier. It is recommended to use keys with business meaning. | region,env,provider |
Tag Value | No | The value corresponding to the key | singapore,production,openai |
Note:
Service Tags are used to automatically filter services within the model API through the "Tag-based routing" policy. When adding a new model service, you only need to configure Tags. The service is then automatically discovered and used by matching APIs, without requiring any modifications to the API configuration.
6. After completing the basic information, click Next to go to the Select Model Policy step.
Model Selection Method: This configuration determines how the gateway handles the model parameter in client requests.
Passthrough Request Model
The gateway will ignore the model parameter in client requests and uniformly use the model you specify in the "Default Model" section below. This mode is suitable for cost control and high availability scenarios, facilitating unified routing and fallback.
Default Model: When the "Model Selection Method" is set to "Specified Model", you must select a specific model name here.
Model Fallback: When it is enabled, the gateway can automatically switch (Fallback) to other available models based on predefined rules if a request to the 'Default Model' fails, ensuring service high availability.
Fallback Rules: After enabling Fallback, you must select or configure the fallback model list and switching rules for when the primary model is unavailable.
The gateway will directly use the model parameter from the client request and forward it to the vendor. This mode is suitable for scenarios that require clients to flexibly control model selection, such as evaluating request latency or counting Token usage. However, when using pass-through requests, the gateway cannot explicitly identify the model name actually requested by the user. Please ensure that the client passes the correct model name.
If the model name in the user request matches the backend vendor's model name or requires no transformation, you can directly pass through the model name from the user request to the backend service.
If you need to replace the model name in the user request with the model name defined by the vendor, you can configure Model Name Mapping information. It supports exact matching and prefix matching (*), and allows you to add multiple mapping rules.
Request Model Name (The model name requested by the client)
Target Model Name (The model name to be rewritten, which is the model name actually used by the backend vendor)
Advanced Configuration (Optional)
Model Parameter Validation: When model parameter validation is enabled, the gateway will validate whether the model parameter in client requests is within the allowed list.
Allowed Model List: Defines the allowlist of model names that clients are allowed to request.
Validation Failure Handling: Defines the handling policy for when model validation fails, supporting "Return 404" or "Fallback to Default Model".
7. After the configuration is completed, click OK to create the model service.
8. After the addition, the newly added service will appear in the service list. Click Service ID/Name to view detailed service information.
Editing a Service
On the Model Service list page, locate the target service and click Edit in its operation column to modify the service configuration information. After making changes, click OK to save.
Deleting a Service
On the Model Service list page, locate the target service and click Delete in its operation column. The system will then perform a dependency check before deletion.
1. The system will display a pop-up window asking you to confirm the deletion and automatically check whether the service is bound to any other resources, such as a model API.
2. Verify results:
If no dependencies exist, the pop-up window will directly display the service ID and name. Click OK to delete it.
If dependencies exist, the pop-up window will display the message "Resource Deletion Dependency Check Result" below the service information, prompt "Unresolved dependencies exist", and list the specific dependency items.
3. If dependencies exist, you must first remove all listed dependencies. After the dependencies are removed, click the Recheck action in the pop-up window. The system will then perform the check again. Once the check passes and the dependency prompt disappears, click OK to finally delete the service. To cancel the deletion, click Cancel.
Model Service Lifecycle Management
Lifecycle State Description
Model services support the following three states. The model service list adds a Status column and an Online/Offline action:
|
Not launched | Manually take offline | No |
Running | Running normally and health check passed. | Yes |
Exception | Health check failed. Waiting for recovery or manual intervention. | No |
State transition trigger conditions:
|
Not launched → Running | Manually bring online | The user clicks "Bring Online" in the console. |
Running → Not launched | Manually take offline | The user clicks "Take Offline" in the console, and the system stops receiving traffic. |
Running → Exception | Health check failure | Consecutive N probe failures (N is configurable, defaulting to 3) |
Exception → Running | Health check recovery | Probe succeeds, and the system recovers automatically (no manual intervention required). |
Exception → Running | Manually bring online | The user confirms that the exception has been resolved and manually triggers the transition. |
Bringing a Model Service Online/Offline
1. In the left sidebar, click Model Management, and then click the Model Service tab.
2. In the model service list, view the Status column of each service.
3. Perform actions on the target service:
Online: Click Online in the operation column. The service enters the "Running" state and starts receiving traffic. Model services are online by default.
Offline: Click Offline in the operation column, and then click Confirm Offline in the confirmation dialog box. After the service is offline, it stops receiving new model API requests, associated model APIs are no longer routed to this service, and the service configuration is retained, allowing the service to be brought back online.
Configuring a Health Check
1. On the model service details page, locate the Health Check Configuration section and click Edit.
2. Enable the health check switch and configure the general parameters:
|
Check Interval | Yes | 30 seconds | Probe interval |
Timeout Time | Yes | 5 (seconds) | Timeout for a single probe |
Failure Threshold | Yes | 3 (times) | Mark as abnormal after N consecutive failures |
Recovery Threshold | Yes | 1 (time) | Mark as recovered after N consecutive successes |
Probe Path | Yes | /v1/models | Configurable based on the model protocol. |