Scenario Overview
This document describes how to quickly integrate TokenHub through the AI Gateway. Tencent Cloud's large model service platform, TokenHub, aggregates model capabilities from multiple model providers such as Hunyuan, DeepSeek, MiniMax, Kimi, Zhipu, and Baidu, and provides services externally through a unified API. After integrating TokenHub through the AI Gateway, you can centrally manage all model services from TokenHub at the gateway layer, gaining unified consumer authentication, quota control, monitoring logs, and link tracing capabilities without integrating each provider's model API individually in your business code.
Applicable scenarios: enterprise-grade AI applications that require unified multi-model access, cost allocation, and traffic control.
Prerequisites
1. An AI Gateway instance has been created, and its status is Running.
2. You have activated the Tencent Cloud TokenHub service and created an API Key on the Quick Access > API Key Management page in the TokenHub console.
Operation Steps
Step 1: Creating a TokenHub Model Service
1. Log in to the AI Gateway console and select the target gateway instance to go to its details page. 2. In the left sidebar, choose Traffic Management > Model Management. Then, select the Model Service tab and click New.
3. In the "Create Model Service" dialog, complete the first step, "Basic Information" configuration:
Scroll down in the "Model Provider" dropdown list (the dropdown uses virtual scrolling), and select TokenHub (Multi-Protocol) under the MaaS Provider category.
Configure the form parameters by referring to the following table:
|
Service Name | Yes | 2-60 characters, supporting uppercase and lowercase letters, digits, and separators ("-", "_"), cannot start with a digit or separator, and cannot end with a separator. |
Service type | - | Fixed to "AI Model Service". The large model capabilities provided by AI Model Service are provided by third parties. Evaluate the service applicability and reliability on your own. |
Model Provider | Yes | Select TokenHub (multi-protocol). |
Model Protocol | Yes | Select OpenAI-compatible (applicable to most models such as DeepSeek, Qwen, and Hunyuan) or Anthropic-compatible (applicable to the Claude model series). |
Service Address | Yes | After a model protocol is selected, the TokenHub domain placeholder is automatically filled in. The actual address is determined by the region. After creation, the actual address of the selected region is displayed. |
Region | Yes | Select the region where the TokenHub service resides (Guangzhou / Singapore), which determines the base domain for calls. After selection, the actual request address is displayed below (such as https://tokenhub.tencentmaas.com/v1/chat/completions) and can be copied with one click. |
Key Credential Type | Yes | Appears after the TokenHub provider is selected. Currently supports API Key. After selection, the model key list below displays only keys of this type. |
Model Key | No (recommended) | Select a created key from the dropdown list (search and refresh supported), or click Create Key to quickly create one. If no key is configured, the gateway cannot carry authentication information when forwarding requests, and calls will return 401. |
Key Usage Policy | - | Default round-robin (load balanced across keys when multiple keys are configured). |
Periodic Key Rotation | No | After selected, the gateway only selects and schedules keys whose creation time falls within the period. |
Description | No | Service description, up to 200 characters. |
4. Click Next to go to the second step, "Select Model Policy":
Model Selection Method:
Specified Model (default): The gateway ignores the model parameter in client requests and uses the model specified in "Default Model" for all requests, which is suitable for cost control and high availability scenarios. You need to select a specific model from the Default Model dropdown list. The dropdown list automatically pulls available models from TokenHub, such as MiniMax-M3, Hy3, GLM-5.3, DeepSeek-V4-Pro, and Kimi K3.
Pass-Through Request Model: The gateway directly forwards the model parameter from client requests to the provider, which is suitable for scenarios where clients need flexible control over model selection.
Model Fallback: After you enable this feature, the gateway can automatically switch to other available models based on rules when a request to the "Default Model" fails, ensuring high service availability.
5. Click OK to complete the model service creation.
Step 2: Viewing Service Details and Credential Management
Click a service name in the model service list to go to the service details page. The details page contains two tabs: Basic Information and Parameter Rewriting.
Basic Information tab:
Basic Information: Service ID, service name, service type, model provider (TokenHub), model protocol, service address (for example, https://tokenhub.tencentmaas.com), region (for example, ap-guangzhou), model selection method, default model, model Fallback, key usage policy, periodic key rotation, creation/modification time, and description. You can click Edit to modify these settings.
Advanced Configuration (collapsible): Service Tag, Timeout Configuration (connection timeout 10000 ms, read timeout 60000 ms, write timeout 60000 ms, number of timeout retries), and Quota Configuration.
Running Status: Displays the current running status (Online) and provides operation buttons for Offline and Configure Health Check.
Credential Management: Displays a list of bound key credentials (ID/Name, Credential, Generation Method, and Actions) and supports search and Add Credential. Make sure that the credential containing the TokenHub API Key is in the "Enabled" state.
Parameter Rewriting tab: You can rewrite request parameters by configuring Parameter Rewriting Rules. If the model name on TokenHub differs from the one used in your business code, you can use a parameter rewriting rule to map the source model name to the target model name without modifying your business code.
Step 3: Creating a Model API
After the model service is created, you need to create a model API and bind it to the service to provide an external access endpoint.
On the Model Management page, select the Model API tab and click New.
Step 1 "Basic Information" is configured as follows:
|
API Name | Yes | 2-60 characters. The naming rules are the same as those for the service name. |
Scenario | - | Select the API purpose. Supported purposes: text generation, text vectorization, image generation, video generation, TTS, tools and metadata, and text ranking. After selection, the system presets a default route based on the scenario. |
Request Protocol | - | Select OpenAI or Anthropic. |
Routing | - | By default, POST /v1/chat/completions is selected (primary route, cannot be deselected); POST /v1/responses is optional (next-generation stateful Agent conversation). |
Base Path | No | Default: /. Set a unified route prefix for this API. |
Description | No | Up to 200 characters. |
Step 2 "Sensitive Information Routing" (optional): After you enable this feature, requests are first processed by sensitive information detection and then routed to specified model services based on the detection results. If you do not need this feature, go to Step 3 directly.
Step 3 "Select Model Service" is configured as follows:
|
Service type | Yes | Select single-model service (fixed routing to one backend service) or multi-model service (distributed by routing policy). |
Tag filtering | No | After it is enabled, model services that meet the conditions are filtered by tag. |
Model Service | Yes | Select a created TokenHub model service from the drop-down list. The system automatically filters matching model services based on the protocol and use case selected in the first step. |
Global cross-service Fallback | No | After it is enabled, global cross-service fallback is enabled. |
Click OK to complete the model API creation.
Step 4: Configuring Authentication and Authorization
Before calling a model API, you need to complete the three-step configuration: "Consumer Key > Consumer > API Authorization".
Create Consumer Key: In the left sidebar, choose the Consumer Management > Consumer Keys tab and click Create:
|
Key Name | Yes | 2-60 characters. The naming rules are the same as those for the service name. |
Key Credential Type | Yes | Supports API Key (default), JWT, OAuth 2.0, and OIDC. |
Generation Method | Yes | Supports custom (default, plaintext storage, suitable for testing and non-sensitive scenarios), KMS (KMS credential), and automatic generation. |
Credential Content | Yes | Entered in custom mode, 8-60 characters, supporting uppercase and lowercase letters, digits, and symbols ("-", "*", "="), where "-" and "*" cannot be used as the first or last character. |
Create Consumer: On the Consumer Management > Consumers tab, click Create:
|
Consumer name | Yes | 2-60 characters. |
Consumer group | No | Select an existing consumer group (such as the automatically created DefaultConsumerGroup). |
Selecting a secret | No | Select the consumer secret created in the previous step. |
Scheduling priority | - | High priority / Medium priority (default) / Low priority. Takes effect when consumer priority scheduling is enabled for the model API. High-priority requests are dequeued first. |
Description | No | Up to 200 characters. |
Authorize the Model API: Go to the model API details page and select the Authentication Policy tab:
Authentication Method: The gateway authenticates requests by using API keys by default. An API can be accessed only by consumers who use the configured authentication method. Ensure that the consumer credential type matches the API authentication method.
Click Add Authorization and select a target in the authorization scope (Authorized Consumer Groups / Authorized Consumers / Authorized Consumer Tags). When you select a consumer group, use the shuttle box to select it (for example, DefaultConsumerGroup), and all consumers in the group automatically gain access.
Step 5: Obtaining the Access Address and Initiating Calls
1. Obtaining the Access Address
Obtain the information in the Basic Information > Network Configuration section of the AI gateway instance:
Private Link: An access address within the VPC (for example, 10.0.1.2), suitable for service calls from the same VPC as the gateway.
Public Network CLB: To enable public network access, click Add to configure a public network CLB and obtain a public VIP or domain name.
Access Port: HTTP 80 / HTTPS 443.
Full access address format: protocol://gateway entry address/route request path, for example, http://10.0.1.2/v1/chat/completions. The Route Management > Access Address section on the model API details page also provides an explanation of this format and request examples (with a copy button).
2. Make a call
curl -X POST "http://{gateway_CLB_address}/v1/chat/completions" \\
-H "Content-Type: application/json" \\
-H "Authorization: Bearer {consumer_key}" \\
-d '{
"model": "MiniMax-M3",
"messages": [
{"role": "user", "content": "Introduce yourself in one sentence."}
],
"max_tokens": 50
}'
3. Parameter description:
{Gateway CLB Address}: The gateway entry address (a private network Private Link address or a public network CLB address), which can be obtained on the Basic Information page of the instance.
{Consumer Key}: The credential content of the consumer key (for example, sk-xxxx), which can be obtained in Consumer Key Management.
The model parameter uses model names from TokenHub Model Plaza, such as MiniMax-M3, DeepSeek-V4-Pro, and Hy3. When the model service is configured as "Specified Model", the gateway ignores this parameter and forwards all requests to the default model.
Authentication method: Add Authorization: Bearer ${your consumer key} to the HTTP request header.
Streaming calls are supported: Add "stream": true to the request body.