tencent cloud

Model list

Download
フォーカスモード
フォントサイズ
最終更新日: 2026-08-28 18:41:54
AI翻訳

Language Models

Model Name
Model (API Parameter)
Supported Capabilities
Context Window
(Tokens)
Maximum Input
(Tokens)
Maximum Output
(Tokens)
Hy4 preview
hy4-preview
Deep Reasoning (Preserved Thinking)
Structured Output
Function Calling
Caching
1M
960k
64k
Hy3
hy3
Deep Reasoning (Preserved Thinking)
Structured Output
Function Calling
Caching
256k
192k
128k
DeepSeek-V4-Flash 0731 GA (Vendor Direct)
deepseek-v4-flash-202605
deepseek/deepseek-v4-flash-0731
deepseek/deepseek-v4-flash
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Pro 0813 GA (Vendor Direct)
deepseek-v4-pro-202606
deepseek/deepseek-v4-pro-0813
deepseek/deepseek-v4-pro
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Flash-Vision-Exp (Vendor Direct)
deepseek/deepseek-v4-flash-vision-exp
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Flash 0731 GA
deepseek-v4-flash-0731
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Pro 0813 GA
deepseek-v4-pro-0813
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Flash
deepseek-v4-flash
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
DeepSeek-V4-Pro
deepseek-v4-pro
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
384k
Deepseek-v3.2
deepseek-v3.2
Deep Reasoning
Structured Output
Function Calling
128k
96k
32k
GLM-5.3
glm-5.3
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
128k
GLM-5.3-Flash
glm-5.3-flash
Deep Reasoning
Structured Output
Function Calling
Caching
Image and Video Understanding
1M
1M
128k
GLM-5.2
glm-5.2
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
128k
GLM-5
glm-5
Deep Reasoning
Function Calling
Caching
200k
200k
128k
GLM-5-Turbo
glm-5-turbo
Deep Reasoning
Structured Output
Function Calling
Caching
200k
200k
128k
GLM-5V-Turbo
glm-5v-turbo
Deep Reasoning
Structured Output
Function Calling
Caching
200k
200k
128k
GLM-5.1
glm-5.1
Deep Reasoning
Structured Output
Function Calling
Caching
200k
200k
128k
Kimi K3
kimi-k3
Deep Reasoning
Structured Output
Function Calling
Caching
1m
1m
1m
Kimi K2.7 Code HighSpeed
kimi-k2.7-code-highspeed
Deep Reasoning
Structured Output
Function Calling
Caching
256k
256k
256k
Kimi K2.7 Code
kimi-k2.7-code
Deep Reasoning
Structured Output
Function Calling
Caching
256k
256k
256k
Kimi-K2.6
kimi-k2.6
Deep Reasoning
Structured Output
Function Calling
Caching
256k
256k
256k
Kimi-K2.5
kimi-k2.5
Deep Reasoning
Structured Output
Function Calling
Caching
256k
224k
16k
MiniMax-M3
minimax-m3
Deep Reasoning
Function Calling
Caching
1M
1M
-
MiniMax-M2.5
minimax-m2.5
Deep Reasoning
Function Calling
Caching
200k
200k
128k
MiniMax-M2.7
minimax-m2.7
Deep Reasoning
Function Calling
Caching
200k
200k
128k
Hy-MT2-Pro
hy-mt2-pro
Hy Translation flagship model, suitable for professional domains and other scenarios that demand high translation quality.
8k
4k
4k
Hy-MT2-Plus
hy-mt2-plus
Translation Model
Leading translation performance with excellent instruction-following capability.
8k
4k
4k
Hy-MT2-Lite
hy-mt2-lite
Hunyuan Translation lightweight model, suitable for scenarios with high requirements for latency.
8k
4k
4k
MiMo-V2.5-Pro
mimo-v2.5-pro
Deep Reasoning
Structured Output
Function Calling
Caching
1M
1M
128k

Vector Models

Model Name
Model (API Parameter)
Model Description
Output Dimension
Context Window (Token)
Kinfra-Text-Embedding-0.6b
kinfra-text-embedding-0.6b
A lightweight text embedding model, suitable for large-scale text retrieval, latency-sensitive, and cost-sensitive scenarios.
1024
32k
Kinfra-Text-Embedding-4b
kinfra-text-embedding-4b
A high-quality text embedding model, suitable for high-quality text search and deep semantic understanding scenarios.
2560
32k
Kinfra-VL-Embedding-2b
kinfra-vl-embedding-2b
A lightweight multimodal embedding model that supports text, image, and video inputs, suitable for multimodal online search, video search, and response-speed-prioritized scenarios.
2048
32k
Kinfra-VL-Embedding-8b
kinfra-vl-embedding-8b
A high-precision multimodal embedding model that supports text, image, and video inputs, suitable for high-precision multimodal search and accuracy-prioritized scenarios.
4096
32k
Note:
All the above models support over 30 mainstream languages, including Chinese, English, Japanese, Korean, French, German, Russian, Portuguese, Spanish, and more.

Vision Models

Image Generation

Model Name
Model (API Parameter)
Model Description
Task Type
Default Concurrency
Hy-Image-3.0
hy-image-v3
Hy Image-3.0 image generation model can think about image layout, composition, and brushwork, and use world knowledge to infer commonsense visuals. It can also parse complex semantics of up to a thousand characters, generate long-form text, complex comics, and memes, and create vivid and engaging educational illustrations.
Text-to-Image
Image-to-Image
5
Seedream-Image-v5.0-pro
Seedream-Image-v5.0-pro
The Seedream-5.0-pro model advances Image Creation to a new stage of controllable production. Its key highlights include more controllable editing, more practical production, and more natural results.
Text-to-Image
Image-to-Image
5
Seedream-Image-v5.0-lite
Seedream-Image-v5.0-lite
A lightweight and fast entry-level version of Seedream 5.0, it is the first to feature real-time web search, combined with precise editing and logical reasoning capabilities. It is suitable for production scenarios involving trending time-sensitive content, complex instructions, and high-concurrency, low-cost requirements.
Text-to-Image
Image-to-Image
5
Vidu-Image-q2
vidu-image-q2
Supports reference-based image generation, text-to-image generation, and image editing, with precise rendering of Chinese and English text and pixel-level restoration of design details such as UI/charts. Suitable for creating posters, infographics, and similar content.
Text-to-Image
Image-to-Image
5

Video Generation

Model Name
Model (API Parameter)
Model Description
Task Type
Default Concurrency
MiniMax-Video-H3
minimax-video-h3
Native multimodal understanding and generation: supports multiple input and output types including text, images, audio, and video, enabling integrated content creation.
Multimodal precise editing and control: supports detail modifications such as replacement and reference, making content adjustments more controllable.
Commercial-grade multi-scenario content generation: covers high-frequency scenarios such as film and television, advertising, gaming, branding, and e-commerce, supporting content production and delivery.
Text-to-Video
Image-to-Video
Reference-to-Video
5
Kling-Video-V3
kling-video-v3
The Kling V3 video generation model supports intelligent storyboarding and 15-second long video generation, delivers up to 4K Ultra quality output, enables scene transitions and continuous storytelling, and is suitable for enterprise advertising and professional film production.
Text-to-Video
Image-to-Video
5
Kling-Video-V3-omni
kling-video-v3-omni
Kling 3.0 Omni is an all-in-one multimodal video model in the Kling 3.0 series. It supports text, image, and video inputs, offers character voice driving, native audio output, and storyboarding capabilities, and is suitable for complex narrative videos, multi-subject consistency, and audio-video synchronized creation.
Text-to-Video
Image-to-Video
Reference-to-Video
5
Kling-Video-V3-turbo
kling-video-v3-turbo
A cost-effective fast edition in the Kling 3.0 series, it supports text and image inputs, improves generation efficiency while maintaining stable output quality, and is suitable for rapid video delivery, marketing clips, creative previews, and cost-sensitive video production scenarios.
Text-to-Video
Image-to-Video
5
PixVerse-Video-v6
pixverse-video-v6.0
Supports cinematic camera control, native audio generation, and continuous multi-shot output, and is suitable for professional film and television production and high-quality advertising production.
Text-to-Video
Image-to-Video
5
PixVerse-Video-c1
pixverse-video-c1
Deeply optimized for the film and television industry, it supports multimodal inputs such as text, images, first and last frames, and multi-grid storyboard references, can directly output 15-second 1080P audio-video synchronized videos, features industrial-grade motion performance and film-grade visual effects rendering capabilities, and is suitable for short dramas, animation, fantasy visual effects, and pre-production storyboarding scenarios.
Text-to-Video
Image-to-Video
5
HY-Video-1.5
hy-video-1.5
Supports text and image multimodal inputs to generate high-definition videos, enables scene transitions and multi-character interactions, simplifies production processes and reduces costs, and applies to enterprise advertising and personal creative implementation scenarios.
Text-to-Video
Image-to-Video
5
The performance and effect statements above are derived from Tencent's internal experimental test results in 2026. The tests were conducted under specific environments, conditions, and time ranges, and only reflect the corresponding test scenarios. Actual results may vary depending on input content, parameter configurations, and usage scenarios.

3D Generation

Model Name
Model (API Parameter)
Model Description
Task Type
Default Concurrency
Hy-World-2.1-panorama
hy-world2-panorama
Hy World Model generates 360-degree panoramas. Input a text description or an image, and the model synthesizes a high-fidelity 360-degree panorama.
Text-to-Panorama
Image-to-Panorama
1
Hy-World-2.1-scene
hy-world2-scene
Hy World Model generates 3D scenes. Input a text description or an image, and the model synthesizes a high-fidelity, walkable 3DGS/Mesh scene. The generated world supports free movement and physical collision, and can be seamlessly integrated into creative engines.
Text-to-3D Scene
Image-to-3D Scene
1

Capability Description

Deep Reasoning

The model, before generating the final response, first performs internal (Chain-of-Thought) reasoning by step-by-step analyzing and decomposing problems, thereby improving the accuracy of responses to complex tasks (such as mathematics, logical reasoning, code generation, and so on).

Structured Output

The model supports outputting structured data in specified formats (such as JSON Schema), facilitating direct parsing and utilization by downstream programs. This capability is suitable for scenarios like information extraction, data population, and API response construction.

Function Calling

The model supports function calling capabilities, which can automatically identify and trigger predefined external tools or APIs during the inference process based on user intent, enabling extended operations such as querying databases and invoking third-party services.

Caching

The model's caching capability can reuse context computation results from historical requests, reducing the overhead of redundant computations, thereby improving response speed and reducing invocation costs.



ヘルプとサポート

この記事はお役に立ちましたか?

フィードバック