Model Name | Model (API Parameter) | Supported Capabilities | Context Window (Tokens) | Maximum Input (Tokens) | Maximum Output (Tokens) |
Hy4 preview | hy4-preview | Deep Reasoning (Preserved Thinking) Structured Output Function Calling Caching | 1M | 960k | 64k |
Hy3 | hy3 | Deep Reasoning (Preserved Thinking) Structured Output Function Calling Caching | 256k | 192k | 128k |
DeepSeek-V4.1-Flash (Vendor Direct) | deepseek/deepseek-flash | Deep Reasoning Structured Output Function Calling Caching Image Understanding | 1M | 1M | 384k |
DeepSeek-V4-Flash 0731 GA (Vendor Direct) | deepseek-v4-flash-202605 deepseek/deepseek-v4-flash-0731 deepseek/deepseek-v4-flash | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Pro 0813 GA (Vendor Direct) | deepseek-v4-pro-202606 deepseek/deepseek-v4-pro-0813 deepseek/deepseek-v4-pro | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Flash-Vision-Exp (Vendor Direct) | deepseek/deepseek-v4-flash-vision-exp | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Flash 0731 GA | deepseek-v4-flash-0731 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Pro 0813 GA | deepseek-v4-pro-0813 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Flash | deepseek-v4-flash | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Pro | deepseek-v4-pro | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
Deepseek-v3.2 | deepseek-v3.2 | Deep Reasoning Structured Output Function Calling | 128k | 96k | 32k |
GLM-5.3 | glm-5.3 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 128k |
GLM-5.3-Flash | glm-5.3-flash | Deep Reasoning Structured Output Function Calling Caching Image and Video Understanding | 1M | 1M | 128k |
GLM-5.2 | glm-5.2 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 128k |
GLM-5 (To be discontinued on 2026-10-08) | glm-5 | Deep Reasoning Function Calling Caching | 200k | 200k | 128k |
GLM-5-Turbo (To be discontinued on 2026-10-08) | glm-5-turbo | Deep Reasoning Structured Output Function Calling Caching | 200k | 200k | 128k |
GLM-5V-Turbo (To be discontinued on 2026-10-30) | glm-5v-turbo | Deep Reasoning Structured Output Function Calling Caching | 200k | 200k | 128k |
GLM-5.1 (To be discontinued on 2026-10-08) | glm-5.1 | Deep Reasoning Structured Output Function Calling Caching | 200k | 200k | 128k |
Kimi K3 | kimi-k3 | Deep Reasoning Structured Output Function Calling Caching | 1m | 1m | 1m |
Kimi K2.8 Preview | kimi-k2.8-preview | Deep Reasoning Structured Output Function Calling Caching | 1m | 1m | 1m |
Kimi K2.7 Code HighSpeed | kimi-k2.7-code-highspeed | Deep Reasoning Structured Output Function Calling Caching | 256k | 256k | 256k |
Kimi K2.7 Code | kimi-k2.7-code | Deep Reasoning Structured Output Function Calling Caching | 256k | 256k | 256k |
Kimi-K2.6 | kimi-k2.6 | Deep Reasoning Structured Output Function Calling Caching | 256k | 256k | 256k |
Kimi-K2.5 | kimi-k2.5 | Deep Reasoning Structured Output Function Calling Caching | 256k | 224k | 16k |
MiniMax-M3 | minimax-m3 | Deep Reasoning Function Calling Caching | 1M | 1M | - |
MiniMax-M2.5 | minimax-m2.5 | Deep Reasoning Function Calling Caching | 200k | 200k | 128k |
MiniMax-M2.7 | minimax-m2.7 | Deep Reasoning Function Calling Caching | 200k | 200k | 128k |
Hy-MT2-Pro | hy-mt2-pro | Hy Translation flagship model, suitable for professional domains and other scenarios that demand high translation quality. | 8k | 4k | 4k |
Hy-MT2-Plus | hy-mt2-plus | Translation Model Leading translation performance with excellent instruction-following capability. | 8k | 4k | 4k |
Hy-MT2-Lite | hy-mt2-lite | Hunyuan Translation lightweight model, suitable for scenarios with high requirements for latency. | 8k | 4k | 4k |
MiMo-V2.5-Pro | mimo-v2.5-pro | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 128k |
Model Name | Model (API Parameter) | Model Description | Output Dimension | Context Window (Token) |
Kinfra-Text-Embedding-0.6b | kinfra-text-embedding-0.6b | A lightweight text embedding model, suitable for large-scale text retrieval, latency-sensitive, and cost-sensitive scenarios. | 1024 | 32k |
Kinfra-Text-Embedding-4b | kinfra-text-embedding-4b | A high-quality text embedding model, suitable for high-quality text search and deep semantic understanding scenarios. | 2560 | 32k |
Kinfra-VL-Embedding-2b | kinfra-vl-embedding-2b | A lightweight multimodal embedding model that supports text, image, and video inputs, suitable for multimodal online search, video search, and response-speed-prioritized scenarios. | 2048 | 32k |
Kinfra-VL-Embedding-8b | kinfra-vl-embedding-8b | A high-precision multimodal embedding model that supports text, image, and video inputs, suitable for high-precision multimodal search and accuracy-prioritized scenarios. | 4096 | 32k |
Model Name | Model (API Parameter) | Model Description | Task Type | Default Concurrency |
Hy-Image-3.0 | hy-image-v3 | Hy Image-3.0 image generation model can think about image layout, composition, and brushwork, and use world knowledge to infer commonsense visuals. It can also parse complex semantics of up to a thousand characters, generate long-form text, complex comics, and memes, and create vivid and engaging educational illustrations. | Text-to-Image Image-to-Image | 5 |
WAND-Vega-Image1.0 Lite | wand-vega-image-lite | WAND-Vega-Image1.0 Lite image generation model offers lower costs and faster generation speeds. It is suitable for large-scale generation scenarios such as e-commerce product images and batch materials, quickly producing usable images within a controlled budget. | Text-to-Image Image-to-Image | 5 |
WAND-Vega-Image1.0 Flash | wand-vega-image-flash | WAND-Vega-Image1.0 Flash image generation model balances generation speed and image quality. It is suitable for daily creation scenarios such as short video covers, social media images, and marketing materials, delivering both efficiency and effectiveness. | Text-to-Image Image-to-Image | 5 |
WAND-Vega-Image1.0 Pro | wand-vega-image-pro | The WAND-Vega-Image1.0 Pro image generation model prioritizes image quality and delivers finer detail. It is suitable for professional creation scenarios such as brand visuals, refined hero images, and high-quality design materials, consistently producing detailed images. | Text-to-Image Image-to-Image | 5 |
Kling-Image-v3 | kling-image-v3 | Uses reference images to precisely edit and modify a single image, combining high image quality with general-purpose creation capabilities. | Text-to-Image Image-to-Image | 5 |
Kling-Image-o1 | kling-image-o1 | With prompts alone, it can perform multiple basic capabilities such as text-to-image and image-to-image generation, focusing on understanding basic text logic and fast generation. | Text-to-Image Image-to-Image | 5 |
Kling-Image-v3-omni | kling-image-v3-omni | It performs better in AI-driven image storytelling and cross-modal fusion, and can deeply optimize image-text generation logic. It is highly suitable for image creation in cinematic or complex visual scenarios. | Text-to-Image Image-to-Image | 5 |
Seedream-Image-v5.0-pro | Seedream-Image-v5.0-pro | The Seedream-5.0-pro model advances Image Creation to a new stage of controllable production. Its key highlights include more controllable editing, more practical production, and more natural results. | Text-to-Image Image-to-Image | 5 |
Seedream-Image-v5.0-lite | Seedream-Image-v5.0-lite | A lightweight and fast entry-level version of Seedream 5.0, it is the first to feature real-time web search, combined with precise editing and logical reasoning capabilities. It is suitable for production scenarios involving trending time-sensitive content, complex instructions, and high-concurrency, low-cost requirements. | Text-to-Image Image-to-Image | 5 |
Vidu-Image-q2 | vidu-image-q2 | Supports reference-based image generation, text-to-image generation, and image editing, with precise rendering of Chinese and English text and pixel-level restoration of design details such as UI/charts. Suitable for creating posters, infographics, and similar content. | Text-to-Image Image-to-Image | 5 |
Model Name | Model (API Parameter) | Model Description | Task Type | Default Concurrency |
MiniMax-Video-H3-Max | minimax-video-h3-max | MiniMax H3 Max, a multimodal video generation model that generates audio and visuals in a single pass, supports first and last frame control and character consistency preservation, and produces 5-second clips in seconds. | Text-to-Video Image-to-Video | 5 |
MiniMax-Video-H3 | minimax-video-h3 | Native multimodal understanding and generation: supports multiple input and output types including text, images, audio, and video, enabling integrated content creation. Multimodal precise editing and control: supports detail modifications such as replacement and reference, making content adjustments more controllable. Commercial-grade multi-scenario content generation: covers high-frequency scenarios such as film and television, advertising, gaming, branding, and e-commerce, supporting content production and delivery. | Text-to-Video Image-to-Video Reference-to-Video | 5 |
Kling-Video-motion-control-v3.0 | kling-video-motion-control-v3.0 | The Kling v2.6 precise motion control and pose transfer model accurately transfers actions, gestures, or facial expressions from a reference video to a static character image. | Motion control | 5 |
Kling-Video-motion-control-v2.6 | kling-video-motion-control-v2.6 | The Kling v3.0 precise motion control and pose transfer model significantly enhances facial feature stability and expression naturalness in complex, multi-angle, and multi-second actions. | Motion control | 5 |
Kling-Video-V3 | kling-video-v3 | The Kling V3 video generation model supports intelligent storyboarding and 15-second long video generation, delivers up to 4K Ultra quality output, enables scene transitions and continuous storytelling, and is suitable for enterprise advertising and professional film production. | Text-to-Video Image-to-Video | 5 |
Kling-Video-V3-omni | kling-video-v3-omni | Kling 3.0 Omni is an all-in-one multimodal video model in the Kling 3.0 series. It supports text, image, and video inputs, offers character voice driving, native audio output, and storyboarding capabilities, and is suitable for complex narrative videos, multi-subject consistency, and audio-video synchronized creation. | Text-to-Video Image-to-Video Reference-to-Video | 5 |
Kling-Video-V3-turbo | kling-video-v3-turbo | A cost-effective fast edition in the Kling 3.0 series, it supports text and image inputs, improves generation efficiency while maintaining stable output quality, and is suitable for rapid video delivery, marketing clips, creative previews, and cost-sensitive video production scenarios. | Text-to-Video Image-to-Video | 5 |
PixVerse-Video-v6 | pixverse-video-v6.0 | Supports cinematic camera control, native audio generation, and continuous multi-shot output, and is suitable for professional film and television production and high-quality advertising production. | Text-to-Video Image-to-Video | 5 |
PixVerse-Video-c1 | pixverse-video-c1 | Deeply optimized for the film and television industry, it supports multimodal inputs such as text, images, first and last frames, and multi-grid storyboard references, can directly output 15-second 1080P audio-video synchronized videos, features industrial-grade motion performance and film-grade visual effects rendering capabilities, and is suitable for short dramas, animation, fantasy visual effects, and pre-production storyboarding scenarios. | Text-to-Video Image-to-Video | 5 |
Hy-Video-v1.5 | hy-video-v1.5 | Supports text and image multimodal inputs to generate high-definition videos, enables scene transitions and multi-character interactions, simplifies production processes and reduces costs, and applies to enterprise advertising and personal creative implementation scenarios. | Text-to-Video Image-to-Video | 5 |
Model Name | Model (API Parameter) | Model Description | Task Type | Default Concurrency |
Hy-World-2.1-panorama | hy-world2-panorama | Hy World Model generates 360-degree panoramas. Input a text description or an image, and the model synthesizes a high-fidelity 360-degree panorama. | Text-to-Panorama Image-to-Panorama | 1 |
Hy-World-2.1-scene | hy-world2-scene | Hy World Model generates 3D scenes. Input a text description or an image, and the model synthesizes a high-fidelity, walkable 3DGS/Mesh scene. The generated world supports free movement and physical collision, and can be seamlessly integrated into creative engines. | Text-to-3D Scene Image-to-3D Scene | 1 |
Model Name | Model (API Parameter) | Model Description | Task Type | Default Concurrency |
Mureka-Music-v9 | mureka-music-v9 | A Kunlun Wanwei music generation model that offers high control and fidelity over musical style, emotion, vocals, and instrumental arrangement. | Music generation | 5 |
IndexTTS-2 | indextts-2 | IndexTTS-2 is an industrial-grade autoregressive zero-shot voice cloning model open-sourced by Bilibili, with controllable emotion and duration. | TTS | 5 |
WAND-Dubbing-STS-v1 | wand-dubbing-sts-v1 | WAND-Dubbing-STS-v1 is a voice timbre conversion model that replaces the speaker's timbre with a specified voice while fully preserving the original speech content: it does not alter the dialogue, retains the original speech rate, pauses, and emotional performance, and only changes the timbre. It is suitable for scenarios such as short video voice changing, character voice unification, and material remastering. | AI dubbing | 5 |
Apakah halaman ini membantu?
Anda juga dapat Menghubungi Penjualan atau Mengirimkan Tiket untuk meminta bantuan.
masukan