tencent cloud

LLM Service TokenHub

Video Generation Model Invocation Overview

Unduh
Mode fokus
Ukuran font
Terakhir diperbarui: 2026-09-10 22:08:17
Diterjemahkan oleh AI

Overview

The TokenHub video generation models provide capabilities such as text-to-video, image-to-video, first-and-last-frame-to-video, reference-to-video, and all-in-one video generation, covering model series including PixVerse, Kling, MiniMax, and Hy. This document serves as a quick navigation and selection guide for each model.
Note:
For detailed parameters, request examples, and response fields of each API, refer to the corresponding document. For billing information, see TokenHub Model Pricing.

Model List

PixVerse Series

For details, see the PixVerse Call Guide.
Model Name
model Parameter Value
Supported Capability
Brief Description
PixVerse-Video-v6
pixverse-video-v6.0
Text-to-video / image-to-video / first-last frame / reference-based generation
Recommended for general scenarios, supports multi-shot intelligent storyboarding.
PixVerse-Video-c1
pixverse-video-c1
Text-to-video / image-to-video / first-last frame / reference-based generation
Recommended for dynamic scenes such as combat, spell effects, and high-speed motion.

Kling Series

For details, see the Kling Call Guide.
Model Name
model Parameter Value
Supported Capability
Brief Description
Kling-Video-V3
kling-video-v3
Text-to-video / image-to-video (including first-last frame and element)
Flagship model: native audio, multi-shot, element reference, and 4K, with the most comprehensive capabilities.
Kling-Video-V3-omni
kling-video-v3-omni
Omnipotent video generation (multimodal input and video editing)
Scenarios with multimodal mixed inputs such as reference video and video editing.
Kling-Video-V3-turbo
kling-video-v3-turbo
Text-to-video / image-to-video
V3 fast version, batch generation and cost-first; no audio or multi-shot.

MiniMax Video

For details, see the MiniMax Call Guide.
Model Name
model Parameter Value
Supported Capability
Brief Description
MiniMax-Video-H3
minimax-video-h3
Text-to-video / image-to-video (first and last frames) / multimodal reference-based generation
Latest flagship model: 2K direct output, multimodal reference, native stereo sound, and up to 15 seconds.
MiniMax-Video-H3-Max
minimax-video-h3-max
Text-to-video / image-to-video (first frame / last frame), middle frames not supported.
A multimodal video generation model that generates audio and visuals in a single pass, supports first and last frame control and character consistency preservation, and produces 5-second clips in seconds.

Hy Video

For details, see the Hy Call Guide.
Model Name
model Parameter Value
Supported Capability
Brief Description
Hy-Video-v1.5
hy-video-v1.5
Text-to-video / image-to-video
Hunyuan video generation intelligently creates videos based on the Hunyuan foundation model.







Bantuan dan Dukungan

Apakah halaman ini membantu?

masukan