Explore APIAny models across chat, image, video, audio, and safety detection — capabilities, request types, pricing, and parameters for every model.
Showing 17 groups

Zhipu GLM-5.2 — a next-gen bilingual (Chinese/English) LLM for general chat, code generation, and agentic tasks, strong in Chinese-language contexts
Input
Output
Input
≈ $1.88-$3.75
Output
≈ $6.25-$12.5
Cache
≈ $0.375-$0.75

Google Gemini 3.5 Flash — a fast, cost-efficient multimodal chat model with image+text input and a million-token context window, for high-concurrency workloads
Input
Output
Input
≈ $0.938-$1.88
Output
≈ $6.25-$12.5
Cache
≈ $0.0938-$0.188

MiniMax M3 — a large language model with a million-token context window, excelling at long-document understanding, complex reasoning, and tool use
Input
Output
Input
≈ $1.2-$2.4
Output
≈ $5-$10
Cache
≈ $1.5-$3

DeepSeek V4 — a reasoning-focused LLM with a 128K-token context window, excelling at code generation and complex logic at a highly competitive price
Input
Output
Input
≈ $0.212-$0.425
Output
≈ $0.475-$0.95
Cache
≈ $0.025-$0.05

OpenAI GPT-5.5 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $4.25-$8.5
Output
≈ $25-$50
Cache
≈ $0.425-$0.85

OpenAI GPT-5.6 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $6-$12
Output
≈ $40-$80
Cache
≈ $0.85-$1.7

Grok 4.5 is an advanced AI model designed to deliver fast, writing, coding, research, data analysis, and creative problem-solving through natural conversation
Input
Output
Input
≈ $0.6-$1.2
Output
≈ $1.2-$2.4
Cache
≈ $0.6-$1.2

A premium reasoning route for visual front-end prototyping, repository-scale coding, large evidence sets, long-running agents, and complex knowledge work that benefits from a 1.05M-token working context.
Input
Output
Input
≈ $2-$4
Output
≈ $10-$20
Cache
≈ $0.2-$0.4

Google Gemini 2.5 Flash Lite — an ultra-low-cost, low-latency chat model with a million-token context window, ideal for high-frequency everyday tasks
Input
Output
Input
≈ $0.1-$0.2
Output
≈ $0.15-$0.3
Cache
≈ $0.01-$0.02

Google Gemini 3.1 Flash Lite — improved reasoning over the 2.5 generation while staying economical, balancing speed and quality for lightweight tasks
Input
Output
Input
≈ $0.15-$0.3
Output
≈ $1-$2
Cache
≈ $0.0125-$0.025

Google Gemini 3.1 Pro — the flagship multimodal Gemini model, offering strong complex reasoning and long-context capability for demanding production apps
Input
Output
Input
≈ $1.25-$2.5
Output
≈ $7.5-$15
Cache
≈ $0.625-$1.25

OpenAI GPT-4o mini — a fast, affordable multimodal chat model with quick responses and low cost, ideal for everyday Q&A and lightweight coding help
Input
Output
Input
≈ $0.275-$0.55
Output
≈ $0.412-$0.825
Cache
≈ $0.0138-$0.0275

OpenAI GPT-5.4 — a high-capability model for advanced reasoning, code generation, and agentic workflows, with a 400K-token context window for production use
Input
Output
Input
≈ $2.5-$5
Output
≈ $15-$30
Cache
≈ $0.25-$0.5

Anthropic Claude Opus 4.8 — the flagship Claude model for the most demanding reasoning, coding, and long-form writing tasks, with excellent long context
Input
Output
Input
≈ $2.5-$5
Output
≈ $12-$24
Cache
≈ $0.25-$0.5

Anthropic Claude Opus 5 from APIAny — Anthropic’s newest Opus-tier flagship for the hardest coding, long-running agents, and judgment-heavy review.
Input
Output
Input
≈ $3.75-$7.5
Output
≈ $18.75-$37.5
Cache
≈ $0.375-$0.75

Anthropic Claude Sonnet 4.6 — a balanced flagship model delivering strong reasoning at production speed and cost, for large-scale, stable deployment
Input
Output
Input
≈ $1.5-$3
Output
≈ $7.5-$15
Cache
≈ $1-$2

Anthropic Claude Sonnet 5 — a balanced flagship model delivering strong reasoning at production speed and cost, for large-scale, stable deployment
Input
Output
Input
≈ $1.88-$3.75
Output
≈ $10-$20
Cache
≈ $1.25-$2.5