tencent cloud

LLM Service TokenHub

Model list

Unduh
Mode fokus
Ukuran font
Terakhir diperbarui: 2026-09-16 17:44:42
Diterjemahkan oleh AI

Language Models

Model Name
Model (API Parameter)
Supported Capabilities
Context Window
(Tokens)
Maximum Input
(Tokens)
Maximum Output
(Tokens)
Hy4 preview
hy4-preview
Deep Reasoning (Preserved Thinking)
Structured Output
Function Calling
Caching
1M
960k
64k
Hy3
hy3
Deep Reasoning (Preserved Thinking)
Structured Output
Function Calling
Caching
256k
192k
128k
DeepSeek-V4.1-Flash (Vendor Direct)
deepseek/deepseek-flash
Deep Reasoning
Structured Output
Function Calling
Caching
Image Understanding
1M
1M
384k
DeepSeek-V4-Flash 0731 GA (Vendor Direct)
deepseek-v4-flash-202605
deepseek/deepseek-v4-flash-0731
deepseek/deepseek-v4-flash
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Pro 0813 GA (Vendor Direct)
deepseek-v4-pro-202606
deepseek/deepseek-v4-pro-0813
deepseek/deepseek-v4-pro
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Flash-Vision-Exp (Vendor Direct)
deepseek/deepseek-v4-flash-vision-exp
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Flash 0731 GA
deepseek-v4-flash-0731
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Pro 0813 GA
deepseek-v4-pro-0813
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Flash
deepseek-v4-flash
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Pro
deepseek-v4-pro
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
Deepseek-v3.2
deepseek-v3.2
Deep Reasoning
Structured Output
Function Calling
128k
96k
32k
GLM-5.3
glm-5.3
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
128k
GLM-5.3-Flash
glm-5.3-flash
Deep Reasoning
Structured Output
Function Calling
Caching
Image and Video Understanding
1M
1M
128k
GLM-5.2
glm-5.2
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
128k
GLM-5
(To be discontinued on 2026-10-08)
glm-5
Deep Reasoning
Function Calling
Caching
200k
200k
128k
GLM-5-Turbo
(To be discontinued on 2026-10-08)
glm-5-turbo
Deep Reasoning
Structured Output
Function Calling
Caching
200k
200k
128k
GLM-5V-Turbo
(To be discontinued on 2026-10-30)
glm-5v-turbo
Deep Reasoning
Structured Output
Function Calling
Caching
200k
200k
128k
GLM-5.1
(To be discontinued on 2026-10-08)
glm-5.1
Deep Reasoning
Structured Output
Function Calling
Caching
200k
200k
128k
Kimi K3
kimi-k3
Deep Reasoning
Structured Output
Function Calling
Caching
1m
1m
1m
Kimi K2.8 Preview
kimi-k2.8-preview
Deep Reasoning
Structured Output
Function Calling
Caching
1m
1m
1m
Kimi K2.7 Code HighSpeed
kimi-k2.7-code-highspeed
Deep Reasoning
Structured Output
Function Calling
Caching
256k
256k
256k
Kimi K2.7 Code
kimi-k2.7-code
Deep Reasoning
Structured Output
Function Calling
Caching
256k
256k
256k
Kimi-K2.6
kimi-k2.6
Deep Reasoning
Structured Output
Function Calling
Caching
256k
256k
256k
Kimi-K2.5
kimi-k2.5
Deep Reasoning
Structured Output
Function Calling
Caching
256k
224k
16k
MiniMax-M3
minimax-m3
Deep Reasoning
Function Calling
Caching
1M
1M
-
MiniMax-M2.5
minimax-m2.5
Deep Reasoning
Function Calling
Caching
200k
200k
128k
MiniMax-M2.7
minimax-m2.7
Deep Reasoning
Function Calling
Caching
200k
200k
128k
Hy-MT2-Pro
hy-mt2-pro
Hy Translation flagship model, suitable for professional domains and other scenarios that demand high translation quality.
8k
4k
4k
Hy-MT2-Plus
hy-mt2-plus
Translation Model
Leading translation performance with excellent instruction-following capability.
8k
4k
4k
Hy-MT2-Lite
hy-mt2-lite
Hunyuan Translation lightweight model, suitable for scenarios with high requirements for latency.
8k
4k
4k
MiMo-V2.5-Pro
mimo-v2.5-pro
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
128k

Vector Models

Model Name
Model (API Parameter)
Model Description
Output Dimension
Context Window (Token)
Kinfra-Text-Embedding-0.6b
kinfra-text-embedding-0.6b
A lightweight text embedding model, suitable for large-scale text retrieval, latency-sensitive, and cost-sensitive scenarios.
1024
32k
Kinfra-Text-Embedding-4b
kinfra-text-embedding-4b
A high-quality text embedding model, suitable for high-quality text search and deep semantic understanding scenarios.
2560
32k
Kinfra-VL-Embedding-2b
kinfra-vl-embedding-2b
A lightweight multimodal embedding model that supports text, image, and video inputs, suitable for multimodal online search, video search, and response-speed-prioritized scenarios.
2048
32k
Kinfra-VL-Embedding-8b
kinfra-vl-embedding-8b
A high-precision multimodal embedding model that supports text, image, and video inputs, suitable for high-precision multimodal search and accuracy-prioritized scenarios.
4096
32k
Note:
All the above models support over 30 mainstream languages, including Chinese, English, Japanese, Korean, French, German, Russian, Portuguese, Spanish, and more.

Vision Models

Image Generation

Model Name
Model (API Parameter)
Model Description
Task Type
Default Concurrency
Hy-Image-3.0
hy-image-v3
Hy Image-3.0 image generation model can think about image layout, composition, and brushwork, and use world knowledge to infer commonsense visuals. It can also parse complex semantics of up to a thousand characters, generate long-form text, complex comics, and memes, and create vivid and engaging educational illustrations.
Text-to-Image
Image-to-Image
5
WAND-Vega-Image1.0 Lite
wand-vega-image-lite
WAND-Vega-Image1.0 Lite image generation model offers lower costs and faster generation speeds. It is suitable for large-scale generation scenarios such as e-commerce product images and batch materials, quickly producing usable images within a controlled budget.
Text-to-Image
Image-to-Image
5
WAND-Vega-Image1.0 Flash
wand-vega-image-flash
WAND-Vega-Image1.0 Flash image generation model balances generation speed and image quality. It is suitable for daily creation scenarios such as short video covers, social media images, and marketing materials, delivering both efficiency and effectiveness.
Text-to-Image
Image-to-Image
5
WAND-Vega-Image1.0 Pro
wand-vega-image-pro
The WAND-Vega-Image1.0 Pro image generation model prioritizes image quality and delivers finer detail. It is suitable for professional creation scenarios such as brand visuals, refined hero images, and high-quality design materials, consistently producing detailed images.
Text-to-Image
Image-to-Image
5
Kling-Image-v3
kling-image-v3
Uses reference images to precisely edit and modify a single image, combining high image quality with general-purpose creation capabilities.
Text-to-Image
Image-to-Image
5
Kling-Image-o1
kling-image-o1
With prompts alone, it can perform multiple basic capabilities such as text-to-image and image-to-image generation, focusing on understanding basic text logic and fast generation.
Text-to-Image
Image-to-Image
5
Kling-Image-v3-omni
kling-image-v3-omni
It performs better in AI-driven image storytelling and cross-modal fusion, and can deeply optimize image-text generation logic. It is highly suitable for image creation in cinematic or complex visual scenarios.
Text-to-Image
Image-to-Image
5
Seedream-Image-v5.0-pro
Seedream-Image-v5.0-pro
The Seedream-5.0-pro model advances Image Creation to a new stage of controllable production. Its key highlights include more controllable editing, more practical production, and more natural results.
Text-to-Image
Image-to-Image
5
Seedream-Image-v5.0-lite
Seedream-Image-v5.0-lite
A lightweight and fast entry-level version of Seedream 5.0, it is the first to feature real-time web search, combined with precise editing and logical reasoning capabilities. It is suitable for production scenarios involving trending time-sensitive content, complex instructions, and high-concurrency, low-cost requirements.
Text-to-Image
Image-to-Image
5
Vidu-Image-q2
vidu-image-q2
Supports reference-based image generation, text-to-image generation, and image editing, with precise rendering of Chinese and English text and pixel-level restoration of design details such as UI/charts. Suitable for creating posters, infographics, and similar content.
Text-to-Image
Image-to-Image
5

Video Generation

Model Name
Model (API Parameter)
Model Description
Task Type
Default Concurrency
MiniMax-Video-H3-Max
minimax-video-h3-max
MiniMax H3 Max, a multimodal video generation model that generates audio and visuals in a single pass, supports first and last frame control and character consistency preservation, and produces 5-second clips in seconds.
Text-to-Video
Image-to-Video
5
MiniMax-Video-H3
minimax-video-h3
Native multimodal understanding and generation: supports multiple input and output types including text, images, audio, and video, enabling integrated content creation.
Multimodal precise editing and control: supports detail modifications such as replacement and reference, making content adjustments more controllable.
Commercial-grade multi-scenario content generation: covers high-frequency scenarios such as film and television, advertising, gaming, branding, and e-commerce, supporting content production and delivery.
Text-to-Video
Image-to-Video
Reference-to-Video
5
Kling-Video-motion-control-v3.0
kling-video-motion-control-v3.0
The Kling v2.6 precise motion control and pose transfer model accurately transfers actions, gestures, or facial expressions from a reference video to a static character image.
Motion control
5
Kling-Video-motion-control-v2.6
kling-video-motion-control-v2.6
The Kling v3.0 precise motion control and pose transfer model significantly enhances facial feature stability and expression naturalness in complex, multi-angle, and multi-second actions.
Motion control
5
Kling-Video-V3
kling-video-v3
The Kling V3 video generation model supports intelligent storyboarding and 15-second long video generation, delivers up to 4K Ultra quality output, enables scene transitions and continuous storytelling, and is suitable for enterprise advertising and professional film production.
Text-to-Video
Image-to-Video
5
Kling-Video-V3-omni
kling-video-v3-omni
Kling 3.0 Omni is an all-in-one multimodal video model in the Kling 3.0 series. It supports text, image, and video inputs, offers character voice driving, native audio output, and storyboarding capabilities, and is suitable for complex narrative videos, multi-subject consistency, and audio-video synchronized creation.
Text-to-Video
Image-to-Video
Reference-to-Video
5
Kling-Video-V3-turbo
kling-video-v3-turbo
A cost-effective fast edition in the Kling 3.0 series, it supports text and image inputs, improves generation efficiency while maintaining stable output quality, and is suitable for rapid video delivery, marketing clips, creative previews, and cost-sensitive video production scenarios.
Text-to-Video
Image-to-Video
5
PixVerse-Video-v6
pixverse-video-v6.0
Supports cinematic camera control, native audio generation, and continuous multi-shot output, and is suitable for professional film and television production and high-quality advertising production.
Text-to-Video
Image-to-Video
5
PixVerse-Video-c1
pixverse-video-c1
Deeply optimized for the film and television industry, it supports multimodal inputs such as text, images, first and last frames, and multi-grid storyboard references, can directly output 15-second 1080P audio-video synchronized videos, features industrial-grade motion performance and film-grade visual effects rendering capabilities, and is suitable for short dramas, animation, fantasy visual effects, and pre-production storyboarding scenarios.
Text-to-Video
Image-to-Video
5
Hy-Video-v1.5
hy-video-v1.5
Supports text and image multimodal inputs to generate high-definition videos, enables scene transitions and multi-character interactions, simplifies production processes and reduces costs, and applies to enterprise advertising and personal creative implementation scenarios.
Text-to-Video
Image-to-Video
5
The performance and effect statements above are derived from Tencent's internal experimental test results in 2026. The tests were conducted under specific environments, conditions, and time ranges, and only reflect the corresponding test scenarios. Actual results may vary depending on input content, parameter configurations, and usage scenarios.

3D Generation

Model Name
Model (API Parameter)
Model Description
Task Type
Default Concurrency
Hy-World-2.1-panorama
hy-world2-panorama
Hy World Model generates 360-degree panoramas. Input a text description or an image, and the model synthesizes a high-fidelity 360-degree panorama.
Text-to-Panorama
Image-to-Panorama
1
Hy-World-2.1-scene
hy-world2-scene
Hy World Model generates 3D scenes. Input a text description or an image, and the model synthesizes a high-fidelity, walkable 3DGS/Mesh scene. The generated world supports free movement and physical collision, and can be seamlessly integrated into creative engines.
Text-to-3D Scene
Image-to-3D Scene
1

Speech Models

Model Name
Model (API Parameter)
Model Description
Task Type
Default Concurrency
Mureka-Music-v9
mureka-music-v9
A Kunlun Wanwei music generation model that offers high control and fidelity over musical style, emotion, vocals, and instrumental arrangement.
Music generation
5
IndexTTS-2
indextts-2
IndexTTS-2 is an industrial-grade autoregressive zero-shot voice cloning model open-sourced by Bilibili, with controllable emotion and duration.
TTS
5
WAND-Dubbing-STS-v1
wand-dubbing-sts-v1
WAND-Dubbing-STS-v1 is a voice timbre conversion model that replaces the speaker's timbre with a specified voice while fully preserving the original speech content: it does not alter the dialogue, retains the original speech rate, pauses, and emotional performance, and only changes the timbre. It is suitable for scenarios such as short video voice changing, character voice unification, and material remastering.
AI dubbing
5

Capability Description

Deep Reasoning

The model, before generating the final response, first performs internal (Chain-of-Thought) reasoning by step-by-step analyzing and decomposing problems, thereby improving the accuracy of responses to complex tasks (such as mathematics, logical reasoning, code generation, and so on).

Structured Output

The model supports outputting structured data in specified formats (such as JSON Schema), facilitating direct parsing and utilization by downstream programs. This capability is suitable for scenarios like information extraction, data population, and API response construction.

Function Calling

The model supports function calling capabilities, which can automatically identify and trigger predefined external tools or APIs during the inference process based on user intent, enabling extended operations such as querying databases and invoking third-party services.

Caching

The model's caching capability can reuse context computation results from historical requests, reducing the overhead of redundant computations, thereby improving response speed and reducing invocation costs.



Bantuan dan Dukungan

Apakah halaman ini membantu?

masukan