Model Name | Model (API Parameter) | Supported Capabilities | Context Window (Tokens) | Maximum Input (Tokens) | Maximum Output (Tokens) |
Hy4 preview | hy4-preview | Deep Reasoning (Preserved Thinking) Structured Output Function Calling Caching | 1M | 960k | 64k |
Hy3 | hy3 | Deep Reasoning (Preserved Thinking) Structured Output Function Calling Caching | 256k | 192k | 128k |
DeepSeek-V4-Flash 0731 GA (Vendor Direct) | deepseek-v4-flash-202605 deepseek/deepseek-v4-flash-0731 deepseek/deepseek-v4-flash | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Pro 0813 GA (Vendor Direct) | deepseek-v4-pro-202606 deepseek/deepseek-v4-pro-0813 deepseek/deepseek-v4-pro | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Flash-Vision-Exp (Vendor Direct) | deepseek/deepseek-v4-flash-vision-exp | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Flash 0731 GA | deepseek-v4-flash-0731 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Pro 0813 GA | deepseek-v4-pro-0813 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Flash | deepseek-v4-flash | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
DeepSeek-V4-Pro | deepseek-v4-pro | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 384k |
Deepseek-v3.2 | deepseek-v3.2 | Deep Reasoning Structured Output Function Calling | 128k | 96k | 32k |
GLM-5.3 | glm-5.3 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 128k |
GLM-5.3-Flash | glm-5.3-flash | Deep Reasoning Structured Output Function Calling Caching Image and Video Understanding | 1M | 1M | 128k |
GLM-5.2 | glm-5.2 | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 128k |
GLM-5 | glm-5 | Deep Reasoning Function Calling Caching | 200k | 200k | 128k |
GLM-5-Turbo | glm-5-turbo | Deep Reasoning Structured Output Function Calling Caching | 200k | 200k | 128k |
GLM-5V-Turbo | glm-5v-turbo | Deep Reasoning Structured Output Function Calling Caching | 200k | 200k | 128k |
GLM-5.1 | glm-5.1 | Deep Reasoning Structured Output Function Calling Caching | 200k | 200k | 128k |
Kimi K3 | kimi-k3 | Deep Reasoning Structured Output Function Calling Caching | 1m | 1m | 1m |
Kimi K2.7 Code HighSpeed | kimi-k2.7-code-highspeed | Deep Reasoning Structured Output Function Calling Caching | 256k | 256k | 256k |
Kimi K2.7 Code | kimi-k2.7-code | Deep Reasoning Structured Output Function Calling Caching | 256k | 256k | 256k |
Kimi-K2.6 | kimi-k2.6 | Deep Reasoning Structured Output Function Calling Caching | 256k | 256k | 256k |
Kimi-K2.5 | kimi-k2.5 | Deep Reasoning Structured Output Function Calling Caching | 256k | 224k | 16k |
MiniMax-M3 | minimax-m3 | Deep Reasoning Function Calling Caching | 1M | 1M | - |
MiniMax-M2.5 | minimax-m2.5 | Deep Reasoning Function Calling Caching | 200k | 200k | 128k |
MiniMax-M2.7 | minimax-m2.7 | Deep Reasoning Function Calling Caching | 200k | 200k | 128k |
Hy-MT2-Pro | hy-mt2-pro | Hy Translation flagship model, suitable for professional domains and other scenarios that demand high translation quality. | 8k | 4k | 4k |
Hy-MT2-Plus | hy-mt2-plus | Translation Model Leading translation performance with excellent instruction-following capability. | 8k | 4k | 4k |
Hy-MT2-Lite | hy-mt2-lite | Hunyuan Translation lightweight model, suitable for scenarios with high requirements for latency. | 8k | 4k | 4k |
MiMo-V2.5-Pro | mimo-v2.5-pro | Deep Reasoning Structured Output Function Calling Caching | 1M | 1M | 128k |
Model Name | Model (API Parameter) | Model Description | Output Dimension | Context Window (Token) |
Kinfra-Text-Embedding-0.6b | kinfra-text-embedding-0.6b | A lightweight text embedding model, suitable for large-scale text retrieval, latency-sensitive, and cost-sensitive scenarios. | 1024 | 32k |
Kinfra-Text-Embedding-4b | kinfra-text-embedding-4b | A high-quality text embedding model, suitable for high-quality text search and deep semantic understanding scenarios. | 2560 | 32k |
Kinfra-VL-Embedding-2b | kinfra-vl-embedding-2b | A lightweight multimodal embedding model that supports text, image, and video inputs, suitable for multimodal online search, video search, and response-speed-prioritized scenarios. | 2048 | 32k |
Kinfra-VL-Embedding-8b | kinfra-vl-embedding-8b | A high-precision multimodal embedding model that supports text, image, and video inputs, suitable for high-precision multimodal search and accuracy-prioritized scenarios. | 4096 | 32k |
Model Name | Model (API Parameter) | Model Description | Task Type | Default Concurrency |
Hy-Image-3.0 | hy-image-v3 | Hy Image-3.0 image generation model can think about image layout, composition, and brushwork, and use world knowledge to infer commonsense visuals. It can also parse complex semantics of up to a thousand characters, generate long-form text, complex comics, and memes, and create vivid and engaging educational illustrations. | Text-to-Image Image-to-Image | 5 |
Seedream-Image-v5.0-pro | Seedream-Image-v5.0-pro | The Seedream-5.0-pro model advances Image Creation to a new stage of controllable production. Its key highlights include more controllable editing, more practical production, and more natural results. | Text-to-Image Image-to-Image | 5 |
Seedream-Image-v5.0-lite | Seedream-Image-v5.0-lite | A lightweight and fast entry-level version of Seedream 5.0, it is the first to feature real-time web search, combined with precise editing and logical reasoning capabilities. It is suitable for production scenarios involving trending time-sensitive content, complex instructions, and high-concurrency, low-cost requirements. | Text-to-Image Image-to-Image | 5 |
Vidu-Image-q2 | vidu-image-q2 | Supports reference-based image generation, text-to-image generation, and image editing, with precise rendering of Chinese and English text and pixel-level restoration of design details such as UI/charts. Suitable for creating posters, infographics, and similar content. | Text-to-Image Image-to-Image | 5 |
Model Name | Model (API Parameter) | Model Description | Task Type | Default Concurrency |
MiniMax-Video-H3 | minimax-video-h3 | Native multimodal understanding and generation: supports multiple input and output types including text, images, audio, and video, enabling integrated content creation. Multimodal precise editing and control: supports detail modifications such as replacement and reference, making content adjustments more controllable. Commercial-grade multi-scenario content generation: covers high-frequency scenarios such as film and television, advertising, gaming, branding, and e-commerce, supporting content production and delivery. | Text-to-Video Image-to-Video Reference-to-Video | 5 |
Kling-Video-V3 | kling-video-v3 | The Kling V3 video generation model supports intelligent storyboarding and 15-second long video generation, delivers up to 4K Ultra quality output, enables scene transitions and continuous storytelling, and is suitable for enterprise advertising and professional film production. | Text-to-Video Image-to-Video | 5 |
Kling-Video-V3-omni | kling-video-v3-omni | Kling 3.0 Omni is an all-in-one multimodal video model in the Kling 3.0 series. It supports text, image, and video inputs, offers character voice driving, native audio output, and storyboarding capabilities, and is suitable for complex narrative videos, multi-subject consistency, and audio-video synchronized creation. | Text-to-Video Image-to-Video Reference-to-Video | 5 |
Kling-Video-V3-turbo | kling-video-v3-turbo | A cost-effective fast edition in the Kling 3.0 series, it supports text and image inputs, improves generation efficiency while maintaining stable output quality, and is suitable for rapid video delivery, marketing clips, creative previews, and cost-sensitive video production scenarios. | Text-to-Video Image-to-Video | 5 |
PixVerse-Video-v6 | pixverse-video-v6.0 | Supports cinematic camera control, native audio generation, and continuous multi-shot output, and is suitable for professional film and television production and high-quality advertising production. | Text-to-Video Image-to-Video | 5 |
PixVerse-Video-c1 | pixverse-video-c1 | Deeply optimized for the film and television industry, it supports multimodal inputs such as text, images, first and last frames, and multi-grid storyboard references, can directly output 15-second 1080P audio-video synchronized videos, features industrial-grade motion performance and film-grade visual effects rendering capabilities, and is suitable for short dramas, animation, fantasy visual effects, and pre-production storyboarding scenarios. | Text-to-Video Image-to-Video | 5 |
HY-Video-1.5 | hy-video-1.5 | Supports text and image multimodal inputs to generate high-definition videos, enables scene transitions and multi-character interactions, simplifies production processes and reduces costs, and applies to enterprise advertising and personal creative implementation scenarios. | Text-to-Video Image-to-Video | 5 |
Model Name | Model (API Parameter) | Model Description | Task Type | Default Concurrency |
Hy-World-2.1-panorama | hy-world2-panorama | Hy World Model generates 360-degree panoramas. Input a text description or an image, and the model synthesizes a high-fidelity 360-degree panorama. | Text-to-Panorama Image-to-Panorama | 1 |
Hy-World-2.1-scene | hy-world2-scene | Hy World Model generates 3D scenes. Input a text description or an image, and the model synthesizes a high-fidelity, walkable 3DGS/Mesh scene. The generated world supports free movement and physical collision, and can be seamlessly integrated into creative engines. | Text-to-3D Scene Image-to-3D Scene | 1 |
フィードバック