OpenAI
gpt-image-2.5-sunburst
OpenAI
Entrée$6.0000/M
Sortie$36.0000/M
Built for premium visual workflows where editing precision and finished-image quality matter most: a better fit for production campaign creative, polished product imagery, and multi-turn edits that need to preserve the subject, composition, and visual treatment. It accepts text and reference-image inputs and produces images, with longer generation times than Flare.
Type d'entrée:
Type de sortie:
Multimodal Output
OpenAI
gpt-image-2.5-flare
OpenAI
Entrée$6.0000/M
Sortie$36.0000/M
A strong default for high-quality image generation and editing: optimized for faster iteration across creator and social content, product experiences, visual search, rapid prototyping, and high-volume generation. It accepts text and reference-image inputs and produces images; choose Sunburst when tighter editing control matters more than speed.
Type d'entrée:
Type de sortie:
Multimodal Output
OpenAI
gpt-6-astra
OpenAI
Contexte1.05M
Entrée$9.5000/M
Sortie$47.5000/M
A flagship end-to-end model for complex reasoning, software engineering, computer use, deep research, and document work. It offers about a 1.05M-token context window, accepts text, images, and documents, returns text, and supports reasoning, search, function calling, structured outputs, code execution, and tool orchestration. Choose it over lower-tier models for long-horizon task quality and execution stability, with higher usage cost as the trade-off.
Type d'entrée:
Type de sortie:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Qwen
qwen3.8-flash
Qwen
Contexte1M
Entrée$0.1492/M
Sortie$0.4476/M
An efficient multimodal workhorse in the Qwen family, with a native million-token context window and strong support for coding assistance, agentic workflows, and visual understanding such as charts. It fits long documents, codebases, high-concurrency applications, and tool-assisted tasks.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
Zhipu
glm-5.3-flash
Zhipu AI
Contexte1.31M
Entrée$0.1190/M
Sortie$0.4200/M
The efficiency-focused Flash model in the GLM-5 family, with native multimodal support and a design aimed at long-context, high-frequency execution. It is a strong fit for coding, agentic workflows, and complex tasks that require image-and-text understanding, especially when production cost matters.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
DeepSeek
deepseek-v4-flash-vision-exp
DeepSeek
Contexte1.05M
Entrée$0.1490/M
Sortie$0.2980/M
An experimental multimodal variant of DeepSeek-V4-Flash for image understanding alongside text. A practical choice for visual Q&A, screenshot and document analysis, while its experimental status makes it better suited to evaluation and flexible workflows than strict production-critical paths.
Type d'entrée:
Type de sortie:
ReasoningTool UseLong Context
Qwen
wan3.0-video-prime
Qwen
PAR SEC$0.0675/s
The speed-first variant of Alibaba's Wan 3.0 line (early access). It shortens generation latency while retaining high-quality video output, making it a good fit for low-latency previews, batch drafts and fast-iterating short-video workflows. If ultimate visual control matters more than speed, the standard version is a better choice.
Type d'entrée:
Type de sortie:
Multimodal Output
Z.ai
coding-glm-5.3-free
free
Contexte1M
Entrée$0.0000/M
Sortie$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
Zhipu
glm-5.3
Zhipu AI
Contexte1M
Entrée$1.1900/M
Sortie$4.1500/M
A flagship workhorse for complex software engineering and long-horizon agent tasks. It uses the same base model as GLM-5.2, with improvements driven by post-training, delivering a 50% gain on Z.ai Code Bench alongside stronger terminal-operation and vulnerability-discovery capabilities. It offers a 1M-token context window, up to 128K output, always-on reasoning with low, high, and max effort levels, function calling, and structured outputs. It currently accepts text only and is not intended for tasks requiring visual understanding.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
Gemini
gemini-3.7-flash-free
free
Contexte1.05M
Entrée$0.0000/M
Sortie$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Type d'entrée:
Type de sortie:
ReasoningTool UseLong Context
Grok
grok-4.6
xAI
Contexte500K
Entrée$0.8400/M
Sortie$2.5200/M
A flagship general-purpose model for high-quality coding, agentic execution, and knowledge work. It fits tasks that need sustained planning, tool collaboration, and complex problem solving; for simple calls where latency or cost matters most, choose a lighter model.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
Qwen
wan3.0-video
Qwen
PAR SEC$0.0450/s
The standard video generation model of Alibaba's Wan 3.0 line (early access). It supports text-to-video, image-to-video and multi-modal reference inputs, generating clips up to roughly 30 seconds with synchronized audio in a single pass — suitable for ads, short drama and narrative content. Official release date not yet announced.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
qwen3.8-max
Qwen
Contexte1M
Entrée$1.7750/M
Sortie$5.3300/M
Qwen's most capable flagship multimodal reasoning model, built for long-horizon coding, professional work, and multi-stage agent tasks. It accepts text, images, and video, offers a 1M-token context window with up to 128K output, and supports function calling, structured outputs, web search, and a code interpreter. It is best used as a high-quality workhorse for complex tasks rather than for the lowest-cost, lowest-latency batch workloads.
Type d'entrée:
Type de sortie:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Doubao
dreamina-seedance-2-5-filter-off
Doubao
Entrée$3.8520/M
Sortie$3.8520/M
A strong fit for turning text or multimodal references into short videos, with 30-second single-pass generation, multi-round extension, flexible referencing, and precise editing as its main strengths. Useful for ads, short films, and connected-shot drafts, while final cuts still benefit from human selection and editing.
Type d'entrée:
Type de sortie:
Multimodal Output
MoonshotAI
kimi-k3
Moonshot
Contexte1.05M
Entrée$3.0000/M
Sortie$15.0000/M
Moonshot flagship model with native vision and up to a 1M-token context window. It currently runs at max reasoning effort and suits long-horizon coding, knowledge work, and complex tasks that coordinate terminal tools. For stable use, preserve the full reasoning history and start a new session in a compatible agent harness.
Type d'entrée:
Type de sortie:
ReasoningTool UseLong Context
MoonshotAI
coding-kimi-k3-free
free
Contexte1.05M
Entrée$0.0000/M
Sortie$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Type d'entrée:
Type de sortie:
ReasoningTool UseLong Context
OpenAI
gpt-5.6-terra
OpenAI
Contexte1.05M
Entrée$2.5000/M
Sortie$15.0000/M
The balanced workhorse tier of GPT-5.6, combining capability and cost. Well suited to production work involving image understanding, tool use, and large source material; choose Sol for the hardest quality-first work or Luna for cost-controlled high volume.
Type d'entrée:
Type de sortie:
ReasoningFunction CallingStructured OutputLong Context
OpenAI
gpt-5.6-sol
OpenAI
Contexte1.05M
Entrée$5.0000/M
Sortie$30.0000/M
The flagship GPT-5.6 tier for complex professional work and quality-first delivery. It supports image input, function calling, structured outputs, and an exceptionally large context window; consider Terra or Luna when latency or unit cost is the primary constraint.
Type d'entrée:
Type de sortie:
ReasoningFunction CallingStructured OutputLong Context
OpenAI
gpt-5.6-luna
OpenAI
Contexte1.05M
Entrée$1.0000/M
Sortie$6.0000/M
The cost-sensitive, high-volume tier of GPT-5.6. A good fit for summarization, classification, rewriting, and batch automation; choose the Sol sibling for complex professional judgment or quality-first multi-step work.
Type d'entrée:
Type de sortie:
ReasoningFunction CallingStructured OutputLong Context
Doubao
doubao-seedream-5-0-pro
Doubao
PAR IMG$0.0450/image
A multimodal image-generation model for professional visual creation. It can generate images from text and perform precise edits with reference images and spatial annotations, with strengths in dense infographics, typography, realistic lighting, and multilingual text rendering.
Type d'entrée:
Type de sortie:
ReasoningMultimodal Output
Grok
grok-4.5
xAI
Contexte500K
Entrée$2.1000/M
Sortie$6.3000/M
A frontier reasoning model for coding, knowledge work, and STEM tasks. It supports image input, function calling, and structured outputs; its 500K-token context suits cross-file analysis and long source material. Prefer a lower-cost model when frontier reasoning is unnecessary.
Type d'entrée:
Type de sortie:
ReasoningFunction CallingStructured OutputLong Context
Claude
claude-fable-5
Anthropic
Contexte1M
Entrée$12.5000/M
Sortie$50.0000/M
Anthropic's most capable widely released model, built for the most demanding reasoning and long-horizon agentic work. Offers a 1M-token context window with 128k max output and always-on adaptive thinking, excelling at autonomously exploring underspecified tasks, planning, and carrying long-running coding and multi-agent orchestration further before it needs human input.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
Claude
claude-sonnet-5
Anthropic
Contexte1M
Entrée$2.0000/M
Sortie$10.0000/M
A daily-driver Sonnet-tier model aimed at bringing near-frontier agentic, coding, and knowledge-work capability at a lower operating cost. Its default 1M-token context and adaptive thinking fit long documents, codebases, and multi-step tool workflows; for the deepest reasoning or restricted high-risk security work, evaluate higher-tier or specialized models.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
Gemini
nano-banana-2-lite
Google
PAR APPEL$0.0200/call
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Type d'entrée:
Type de sortie:
Multimodal Output
LongCat
LongCat-2.0
美团
Contexte1M
Entrée$0.3000/M
Sortie$1.2000/M
A LongCat 2.0 workhorse for project-scale coding and long-running agent tasks, with native 1M-token context plus tool calling and multi-step reasoning. It is best when you need to keep large repositories, long documents, or automation workflows in scope; for lightweight Q&A or low-latency chat, a smaller Flash/Lite-style model is usually cheaper.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingLong Context
Doubao
doubao-seed-2-1-pro
Doubao
Contexte256K
Entrée$0.9700/M
Sortie$4.8800/M
Built for coding, agents, and complex productivity tasks, advancing beyond Doubao Seed 2.0 in long context, long output, and tool-oriented workflows. Compared with Qwen, GLM, and DeepSeek peers, it is a cost-effective option for Chinese office work, coding, and multi-step automation.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
Qwen
happyhorse-1.1-t2v
阿里巴巴
PAR SEC$0.1350/s
Built for text-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
happyhorse-1.1-r2v
阿里巴巴
PAR SEC$0.1350/s
Built for reference-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
happyhorse-1.1-i2v
阿里巴巴
PAR SEC$0.1350/s
Built for image-to-video, with a stronger audio-native workflow than HappyHorse 1.0, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Type d'entrée:
Type de sortie:
Multimodal Output
Doubao
doubao-seedance-2-0-mini
Doubao
Entrée$0.6300/M
Sortie$0.6300/M
Built for video generation and multi-asset video creation, improving on Seedance 1.x with more complete motion stability, audiovisual generation, and reference inputs. Compared with Veo, Kling, and Runway, it is better suited to Chinese creative workflows driven by mixed text, image, audio, and video inputs. The Mini tier is positioned for lower-cost, high-frequency testing compared with standard Seedance 2.0. Useful for short drama, ads, e-commerce, and social assets.
Type d'entrée:
Type de sortie:
Multimodal Output
Zhipu
glm-5.2
Zhipu AI
Contexte1M
Entrée$1.1900/M
Sortie$4.1400/M
Built for reasoning, coding, and long-context agent tasks, improving over GLM-5.1 in context length, tool use, and complex multi-step workflows. Compared with Chinese flagship peers such as Qwen, DeepSeek, and Kimi, it fits enterprise automation, project-level code analysis, and knowledge work in Chinese.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
MoonshotAI
kimi-k2.7-code-highspeed
Moonshot
Contexte256K
Entrée$1.9000/M
Sortie$7.9999/M
Built for long-context, multi-step tool use, and code/document workflows, with stronger engineering-task stability and context capacity than Kimi K2.6. Compared with Qwen, GLM, and DeepSeek, it is well suited to Chinese long-document analysis, project Q&A, and everyday agentic programming.
Type d'entrée:
Type de sortie:
Tool UseFunction CallingStructured OutputLong Context
MoonshotAI
kimi-k2.7-code
Moonshot
Contexte256K
Entrée$0.9500/M
Sortie$4.0000/M
Built for long-context, multi-step tool use, and code/document workflows, with stronger engineering-task stability and context capacity than Kimi K2.6. Compared with Qwen, GLM, and DeepSeek, it is well suited to Chinese long-document analysis, project Q&A, and everyday agentic programming.
Type d'entrée:
Type de sortie:
Tool UseFunction CallingStructured OutputLong Context
Grok
grok-imagine-video-1-5-preview
xAI
PAR SEC$0.1400/s
Built for image/video-driven generation, editing, and extension, with stronger motion consistency and fast social-content production than Grok Imagine 1.0. Compared with Veo, Kling, and Runway, it fits short-form videos, ad concepts, and trend-driven assets in the xAI ecosystem.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
qwen3.7-plus
Qwen
Contexte1M
Entrée$0.2950/M
Sortie$1.1900/M
Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
Minimax
minimax-m3-free
free
Contexte1M
Entrée$0.0000/M
Sortie$0.0000/M
Free models are for learning and trial purposes only and may be removed at any time. Premium models have limited quotas with 5 requests per minute, and their availability is not guaranteed.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
Minimax
MiniMax-M3
MiniMax
Contexte1M
Entrée$0.3150/M
Sortie$1.2600/M
Built for long-horizon coding, tool use, and multi-turn production collaboration, extending beyond MiniMax M2.7 with multimodal inputs while keeping a 1M-token context window. Compared with Claude, Gemini, and GLM-5.2 long-context agent models, it is better for putting text, image, and video materials into one workflow.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
Claude
claude-opus-4-8
Anthropic
Contexte1M
Entrée$5.0000/M
Sortie$25.0000/M
Anthropic current Opus flagship for the most demanding reasoning, long-running agents, and high-autonomy coding. It supports a 1M-token context window and adaptive thinking; choose it when quality and reliability matter more than latency or cost.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
Gemini
gemini-omni-video
Google
PAR SEC$0.2450/s
Built for any-input-to-video multimodal creation, putting stronger emphasis than earlier Gemini video workflows on unified orchestration of text, image, audio, and video materials. Compared with Veo, Kling, and Runway, it is better for integrating complex multimodal assets into one creative workflow.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
qwen3.7-max
Qwen
Contexte1M
Entrée$1.7744/M
Sortie$5.3197/M
Designed for agentic workloads, coding, and office productivity, with a stronger combined focus on long context, tool calling, and structured output than Qwen3.6. Compared with flagship models such as GPT, Claude, and Gemini, it is a cost-effective alternative for Chinese, coding, and multi-step tool workflows.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
Qwen
happyhorse-1.0-video-edit
Qwen
Contexte8K
PAR SEC$0.1350/s
Built for video editing, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
happyhorse-1.0-t2v
Qwen
Contexte8K
PAR SEC$0.1350/s
Built for text-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
happyhorse-1.0-r2v
Qwen
Contexte8K
PAR SEC$0.1350/s
Built for reference-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Type d'entrée:
Type de sortie:
Multimodal Output
Qwen
happyhorse-1.0-i2v
Qwen
Contexte8K
PAR SEC$0.1350/s
Built for image-to-video, with a stronger audio-native workflow than 早期视频模型, generating dialogue, sound effects, and background music in one pass. Compared with Veo, Kling, and Runway, its differentiator is synchronized audiovisual generation and character-consistent workflows for short drama, ads, e-commerce, and brand marketing.
Type d'entrée:
Type de sortie:
Multimodal Output
DeepSeek
deepseek-v4-pro
DeepSeek
Contexte1M
Entrée$1.7800/M
Sortie$3.5500/M
Built for reasoning, coding, and agentic tool use, with stronger long-context, thinking-mode, and complex-task execution than DeepSeek V3.2. Compared with Chinese peers such as Qwen, GLM, and Kimi, it is especially competitive for math, programming, and logical reasoning tasks.
Type d'entrée:
Type de sortie:
Long ContextTool UseReasoningStructured Output
DeepSeek
deepseek-v4-flash
DeepSeek
Contexte1M
Entrée$0.1800/M
Sortie$0.3500/M
Built for reasoning, coding, and agentic tool use, with stronger long-context, thinking-mode, and complex-task execution than DeepSeek V3.2. Compared with Chinese peers such as Qwen, GLM, and Kimi, it is especially competitive for math, programming, and logical reasoning tasks.
Type d'entrée:
Type de sortie:
Long ContextTool UseReasoningStructured Output
OpenAI
gpt-5.5
OpenAI
Contexte1M
Entrée$5.0000/M
Sortie$30.0000/M
Built for reasoning, coding, visual understanding, and agent tasks, with stronger long context, tool use, and scalable reliable output than GPT-5.4. Compared with Claude, Gemini, and Qwen flagship models, it is a strong general-purpose base for reliable automation and complex knowledge workflows.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
XiaomiMiMo
xiaomi-mimo-v2.5-pro
XiaomiMiMo
Contexte1M
Entrée$75.0000/M
Sortie$75.0000/M
Built for agents, coding, and long-context tasks, offering more complete multimodal understanding, a million-token context window, and long-horizon reasoning than MiMo V2. Compared with Chinese peers such as Qwen, GLM, and DeepSeek, its strengths are Xiaomi ecosystem fit and engineering automation scenarios.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
XiaomiMiMo
xiaomi-mimo-v2.5
XiaomiMiMo
Contexte1M
Entrée$75.0000/M
Sortie$75.0000/M
Built for agents, coding, and long-context tasks, offering more complete multimodal understanding, a million-token context window, and long-horizon reasoning than MiMo V2. Compared with Chinese peers such as Qwen, GLM, and DeepSeek, its strengths are Xiaomi ecosystem fit and engineering automation scenarios.
Type d'entrée:
Type de sortie:
ReasoningTool UseFunction CallingStructured OutputLong Context
XiaomiMiMo
mimo-v2.5-tts-voicedesign
XiaomiMiMo
Entrée$0.0000/M
Sortie$0.0000/M
A voice-design TTS model that creates a new voice from a single natural-language description, useful for brand IP, game NPCs, virtual creators, audio drama characters, and other original-voice workflows.
Type d'entrée:
Type de sortie:
语音合成音色设计风格指令音频标签控制
XiaomiMiMo
mimo-v2.5-tts-voiceclone
XiaomiMiMo
Entrée$0.0000/M
Sortie$0.0000/M
A MiMo v2.5 TTS model focused on voice cloning. It can generate speech close to a target timbre from reference audio while using text instructions to guide style, tone, and audio tags. Best for character dubbing, personalized narration, brand voices, and reusable voice assets.
Type d'entrée:
Type de sortie:
语音合成音色克隆风格指令音频标签控制
XiaomiMiMo
mimo-v2.5-tts
XiaomiMiMo
Entrée$0.0000/M
Sortie$0.0000/M
An out-of-the-box text-to-speech model with multiple premium built-in voices and fine-grained control over pace, emotion, and tone, suited for narration, podcasts, character dubbing, and multi-scenario audio generation.
Type d'entrée:
Type de sortie:
语音合成音频输出风格指令音频标签控制
XiaomiMiMo
mimo-v2.5-pro
XiaomiMiMo
Contexte1M
Entrée$1.1700/M
Sortie$3.5000/M
Built for agents, coding, and long-context tasks, offering more complete multimodal understanding, a million-token context window, and long-horizon reasoning than MiMo V2. Compared with Chinese peers such as Qwen, GLM, and DeepSeek, its strengths are Xiaomi ecosystem fit and engineering automation scenarios.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
XiaomiMiMo
mimo-v2.5
XiaomiMiMo
Contexte1M
Entrée$0.4700/M
Sortie$2.3600/M
Built for agents, coding, and long-context tasks, offering more complete multimodal understanding, a million-token context window, and long-horizon reasoning than MiMo V2. Compared with Chinese peers such as Qwen, GLM, and DeepSeek, its strengths are Xiaomi ecosystem fit and engineering automation scenarios.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
OpenAI
gpt-image-2-text-to-image
OpenAI
Contexte128K
PAR IMG$0.0200/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Type d'entrée:
Type de sortie:
Multimodal Output
OpenAI
gpt-image-2-image-to-image
OpenAI
Contexte128K
PAR IMG$0.0200/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Type d'entrée:
Type de sortie:
Multimodal Output
OpenAI
gpt-image-2
OpenAI
Contexte128K
PAR IMG$0.0670/image
Built for high-quality image generation and editing, improving over GPT Image 1.x in text rendering, local editing, and multi-turn consistency. Compared with Gemini image models, Qwen Image, and Seedream, it is well suited to brand visuals, ad design, product imagery, and tightly controlled creative workflows.
Type d'entrée:
Type de sortie:
Multimodal Output
Claude
claude-opus-4-7
Anthropic
Contexte1M
Entrée$5.0000/M
Sortie$25.0000/M
A high-capability Opus model for complex, long-running software engineering and agent tasks. Compared with Opus 4.6, it emphasizes precise instruction following, higher-resolution vision, sustained execution, and self-verification for large codebases, difficult debugging, document analysis, and multi-step tool use.
Type d'entrée:
Type de sortie:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Minimax
MiniMax-M2.7
MiniMax
Contexte200K
Entrée$0.3100/M
Sortie$1.2400/M
Built for text generation, reasoning, and tool use, with stronger 200K long-context and agentic workflow stability than earlier MiniMax M2.x versions. Compared with Chinese peers such as Kimi, GLM, and Qwen, it fits productivity, coding assistance, and long-document processing.
Type d'entrée:
Type de sortie:
Reasoning
OpenAI
gpt-5.4-nano
OpenAI
Contexte400K
Entrée$0.2000/M
Sortie$1.2500/M
Built for reasoning, coding, visual understanding, and agent tasks, with stronger long context, tool use, and scalable reliable output than GPT-5.3. Compared with Claude, Gemini, and Qwen flagship models, it is a strong general-purpose base for reliable automation and complex knowledge workflows.
Type d'entrée:
Type de sortie:
Tool UseStructured OutputLong Context
OpenAI
gpt-5.4-mini
OpenAI
Contexte400K
Entrée$0.7500/M
Sortie$4.5000/M
Built for reasoning, coding, visual understanding, and agent tasks, with stronger long context, tool use, and scalable reliable output than GPT-5.3. Compared with Claude, Gemini, and Qwen flagship models, it is a strong general-purpose base for reliable automation and complex knowledge workflows.
Type d'entrée:
Type de sortie:
Long ContextTool UseReasoningStructured Output
OpenAI
gpt-5.4
OpenAI
Contexte1.05M
Entrée$2.5000/M
Sortie$15.0000/M
Built for reasoning, coding, visual understanding, and agent tasks, with stronger long context, tool use, and scalable reliable output than GPT-5.3. Compared with Claude, Gemini, and Qwen flagship models, it is a strong general-purpose base for reliable automation and complex knowledge workflows.
Type d'entrée:
Type de sortie:
Long ContextTool UseReasoningStructured Output
Qwen
qwen-image-2.0-pro
Qwen
Contexte33K
Entrée$2.2000/M
Sortie$2.2000/M
Built for image generation and editing, improving over Qwen Image 1.x in Chinese text rendering, instruction following, and complex layouts. Compared with GPT Image, Gemini image models, and Seedream, it is well suited to Chinese posters, e-commerce images, infographics, and iterative visual creation.
Type d'entrée:
Type de sortie:
Structured OutputMultimodal Output
Qwen
qwen-image-2.0
Qwen
Contexte33K
Entrée$2.2000/M
Sortie$2.2000/M
Built for image generation and editing, improving over Qwen Image 1.x in Chinese text rendering, instruction following, and complex layouts. Compared with GPT Image, Gemini image models, and Seedream, it is well suited to Chinese posters, e-commerce images, infographics, and iterative visual creation.
Type d'entrée:
Type de sortie:
Multimodal Output
Gemini
gemini-3.1-flash-image-preview
Google
Contexte131K
PAR IMG$0.0672/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Type d'entrée:
Type de sortie:
ReasoningMultimodal Output
Gemini
nano-banana-2
Google
PAR IMG$0.0330/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Type d'entrée:
Type de sortie:
Multimodal Output
Claude
claude-sonnet-4-6
Anthropic
Contexte1M
Entrée$3.0000/M
Sortie$15.0000/M
The Claude lineup workhorse, balancing strong intelligence, response speed, and cost. It supports a 1M-token context window and adaptive thinking, making it a strong starting point for coding, document analysis, data work, and tool-using agents; move to Opus for the most complex tasks.
Type d'entrée:
Type de sortie:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Doubao
doubao-seed-2-0-pro
Doubao
Contexte256K
Entrée$0.5000/M
Sortie$2.5300/M
Built for coding, agents, and complex productivity tasks, advancing beyond Doubao Seed 2.0 in long context, long output, and tool-oriented workflows. Compared with Qwen, GLM, and DeepSeek peers, it is a cost-effective option for Chinese office work, coding, and multi-step automation.
Type d'entrée:
Type de sortie:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong Context
Doubao
doubao-seedream-5.0-lite
Doubao
PAR IMG$0.0350/image
Built for image generation and editing, with stronger Chinese instruction following, visual reasoning, and high-frequency creative production than Seedream 4.x. Compared with GPT Image, Gemini image models, and Qwen Image, it is well suited to ByteDance-ecosystem ad design, e-commerce assets, and multi-scenario visual production.
Type d'entrée:
Type de sortie:
Web SearchMultimodal Output
Claude
claude-opus-4-6
Anthropic
Contexte1M
Entrée$5.0000/M
Sortie$25.0000/M
An Opus model for demanding reasoning, agentic coding, and large-codebase work. Compared with Opus 4.5, it emphasizes longer-horizon planning, code review, debugging, autonomous execution, and document and spreadsheet work, with a 1M-token context window and adaptive thinking. Evaluate newer Opus 4.7 or 4.8 first for new projects.
Type d'entrée:
Type de sortie:
ReasoningWeb SearchTool UseFunction CallingStructured OutputLong ContextCode Execution
Kling
kling-video-3.0
kling
PAR SEC$0.0900/s
Built for text/image-to-video and multi-shot creation, improving over Kling 2.x in motion stability, camera language, and reference consistency. Compared with Veo, Runway, and Seedance, it is well suited to Chinese short drama, ad assets, and high-quality social video production.
Type d'entrée:
Type de sortie:
Multimodal Output
Kling
kling-v3-omni
kling
PAR SEC$0.0900/s
Built for text/image-to-video and multi-shot creation, improving over Kling 2.x in motion stability, camera language, and reference consistency. Compared with Veo, Runway, and Seedance, it is well suited to Chinese short drama, ad assets, and high-quality social video production.
Type d'entrée:
Type de sortie:
Multimodal Output
Kling
kling-v3
kling
PAR SEC$0.0900/s
Built for text/image-to-video and multi-shot creation, improving over Kling 2.x in motion stability, camera language, and reference consistency. Compared with Veo, Runway, and Seedance, it is well suited to Chinese short drama, ad assets, and high-quality social video production.
Type d'entrée:
Type de sortie:
Multimodal Output
Zhipu
glm-image
Zhipu AI
Entrée$0.8800/M
Sortie$3.5500/M
Built for image generation and editing, with a stronger focus than earlier GLM vision capabilities on Chinese prompts, image-text understanding, and controllable generation. Compared with GPT Image, Gemini image models, and Qwen Image, it fits Chinese marketing assets, explanatory visuals, and iterative visual edits.
Type d'entrée:
Type de sortie:
Multimodal Output
OpenAI
gpt-image-1.5
OpenAI
Contexte128K
Entrée$8.5600/M
Sortie$34.0003/M
Built for high-quality image generation and editing, improving over GPT Image 1.x in text rendering, local editing, and multi-turn consistency. Compared with Gemini image models, Qwen Image, and Seedream, it is well suited to brand visuals, ad design, product imagery, and tightly controlled creative workflows.
Type d'entrée:
Type de sortie:
Multimodal Output
Gemini
gemini-3-pro-image-preview
Google
Contexte66K
PAR IMG$0.1340/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Type d'entrée:
Type de sortie:
ReasoningMultimodal Output
Gemini
nano-banana-pro
Google
PAR IMG$0.0530/image
This is a discounted trial route. Generation may take longer and has roughly a 10% chance of failing. It is best for cost-sensitive image generation trials where queueing and occasional failures are acceptable; for more stable production use, choose a regular route.
Type d'entrée:
Type de sortie:
Multimodal Output
Claude
claude-opus-4-5
Anthropic
Contexte200K
Entrée$5.0000/M
Sortie$25.0000/M
An earlier Opus model in the Claude 4 family, focused on difficult coding, agents, computer use, deep research, and work with slides and spreadsheets. It fits workflows that require Opus-level capability with 4.5 compatibility; without that constraint, evaluate the newer Opus 4.6 or 4.7 first.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
Gemini
veo3.1-lite
Google
PAR APPEL$0.1000/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion. The Lite tier favors high-throughput, lower-cost drafts over the standard tier.
Type d'entrée:
Type de sortie:
Multimodal Output
Gemini
veo3.1-fast
Google
PAR APPEL$0.2010/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion. The Fast tier is better for rapid concept validation than the standard tier.
Type d'entrée:
Type de sortie:
Multimodal Output
Gemini
veo3.1
Google
PAR APPEL$1.5040/call
Built for high-fidelity video generation, with a more mature native-audio, camera-control, and reference-input workflow than Veo 2/3. Compared with Kling, Runway, and Seedance, it is better suited to cinematic shorts, ad storyboards, and videos requiring stable physical motion.
Type d'entrée:
Type de sortie:
Multimodal Output
Gemini
gemini-2.5-flash-image
Google
Contexte66K
PAR IMG$0.0585/image
Built for image generation and editing, with stronger multimodal understanding, conversational edits, and rapid visual prototyping than earlier Gemini image capabilities. Compared with GPT Image, Qwen Image, and Seedream, it is useful for combining documents, images, and prompts in one visual workflow.
Type d'entrée:
Type de sortie:
Structured OutputMultimodal Output
Claude
claude-haiku-4-5
Anthropic
Contexte200K
Entrée$1.1000/M
Sortie$5.5000/M
The fastest and more cost-efficient model in the Claude lineup, offering near-frontier capability with a 200K-token context window and thinking support. It fits real-time interaction, batch processing, lightweight agents, and cost-sensitive production; use Sonnet or Opus for the hardest reasoning or code changes.
Type d'entrée:
Type de sortie:
ReasoningTool UseStructured OutputLong Context
Jina
jina-embeddings-v4
Jina
Contexte33K
Entrée$0.0374/M
Sortie$0.0374/M
A multimodal embedding model for RAG and search infrastructure, not a chat model. It supports 32K-token text inputs and maps text, images, and visual documents into a shared vector space for knowledge-base retrieval, cross-modal/document search, code retrieval, and semantic matching; dense embeddings default to 2048 dimensions and can be truncated to reduce storage cost.
Type d'entrée:
Type de sortie:
OpenAI
gpt-image-2-5-sunburst-text-to-image
OpenAI
PAR SEC$0.0340/s
OpenAI
gpt-image-2-5-sunburst-image-to-image
OpenAI
PAR SEC$0.0340/s
OpenAI
gpt-image-2-5-flare-text-to-image
OpenAI
PAR SEC$0.0340/s
OpenAI
gpt-image-2-5-flare-image-to-image
OpenAI
PAR SEC$0.0340/s
Doubao
doubao-seedance-2-5
Doubao
Entrée$1.3288/M
Sortie$1.3288/M
Doubao
doubao-seedance-2-0-filter-off
Doubao
Entrée$2.1830/M
Sortie$2.1830/M
Doubao
doubao-seedance-2-0-fast-filter-off
Doubao
Entrée$2.0160/M
Sortie$2.0160/M
Doubao
doubao-seedance-2-0-fast
Doubao
Entrée$1.5465/M
Sortie$1.5465/M
Doubao
doubao-seedance-2-0
Doubao
Entrée$2.1606/M
Sortie$2.1606/M
Doubao
dreamina-seedance-2-0-mini-filter-off
Doubao
Entrée$1.2600/M
Sortie$1.2600/M