Showing 18 groups

Zhipu GLM-5.2 — a next-gen bilingual (Chinese/English) LLM for general chat, code generation, and agentic tasks, strong in Chinese-language contexts
Input
Output
Input
≈ $1.88-$3.75
Output
≈ $6.25-$12.5
Cache
≈ $0.375-$0.75

Google Gemini 3.5 Flash — a fast, cost-efficient multimodal chat model with image+text input and a million-token context window, for high-concurrency workloads
Input
Output
Input
≈ $0.938-$1.88
Output
≈ $6.25-$12.5
Cache
≈ $0.0938-$0.188

MiniMax M3 — a large language model with a million-token context window, excelling at long-document understanding, complex reasoning, and tool use
Input
Output
Input
≈ $1.2-$2.4
Output
≈ $5-$10
Cache
≈ $1.5-$3

DeepSeek V4 — a reasoning-focused LLM with a 128K-token context window, excelling at code generation and complex logic at a highly competitive price
Input
Output
Input
≈ $0.212-$0.425
Output
≈ $0.475-$0.95
Cache
≈ $0.025-$0.05

OpenAI GPT-5.5 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $4.25-$8.5
Output
≈ $25-$50
Cache
≈ $0.425-$0.85

OpenAI GPT-5.6 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $6-$12
Output
≈ $40-$80
Cache
≈ $0.85-$1.7
OpenAI GPT-6 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $3.5-$7
Output
≈ $30-$60
Cache
≈ $0.531-$1.06

Grok 4.5 is an advanced AI model designed to deliver fast, writing, coding, research, data analysis, and creative problem-solving through natural conversation
Input
Output
Input
≈ $0.6-$1.2
Output
≈ $1.2-$2.4
Cache
≈ $0.6-$1.2
Grok 4.6 is an advanced AI model designed to deliver fast, writing, coding, research, data analysis, and creative problem-solving through natural conversation
Input
Output
Input
≈ $0.6-$1.2
Output
≈ $1.2-$2.4
Cache
≈ $0.6-$1.2

A premium reasoning route for visual front-end prototyping, repository-scale coding, large evidence sets, long-running agents, and complex knowledge work that benefits from a 1.05M-token working context.
Input
Output
Input
≈ $2-$4
Output
≈ $10-$20
Cache
≈ $0.2-$0.4

Google Gemini 2.5 Flash Lite — an ultra-low-cost, low-latency chat model with a million-token context window, ideal for high-frequency everyday tasks
Input
Output
Input
≈ $0.1-$0.2
Output
≈ $0.15-$0.3
Cache
≈ $0.01-$0.02

Google Gemini 3.1 Flash Lite — improved reasoning over the 2.5 generation while staying economical, balancing speed and quality for lightweight tasks
Input
Output
Input
≈ $0.15-$0.3
Output
≈ $1-$2
Cache
≈ $0.0125-$0.025

Google Gemini 3.1 Pro — the flagship multimodal Gemini model, offering strong complex reasoning and long-context capability for demanding production apps
Input
Output
Input
≈ $1.25-$2.5
Output
≈ $7.5-$15
Cache
≈ $0.625-$1.25

OpenAI GPT-4o mini — a fast, affordable multimodal chat model with quick responses and low cost, ideal for everyday Q&A and lightweight coding help
Input
Output
Input
≈ $0.275-$0.55
Output
≈ $0.412-$0.825
Cache
≈ $0.0138-$0.0275

OpenAI GPT-5.4 — a high-capability model for advanced reasoning, code generation, and agentic workflows, with a 400K-token context window for production use
Input
Output
Input
≈ $2.5-$5
Output
≈ $15-$30
Cache
≈ $0.25-$0.5

Anthropic Claude Opus 4.8 — the flagship Claude model for the most demanding reasoning, coding, and long-form writing tasks, with excellent long context
Input
Output
Input
≈ $2.5-$5
Output
≈ $12-$24
Cache
≈ $0.25-$0.5

Anthropic Claude Opus 5 from APIAny — Anthropic’s newest Opus-tier flagship for the hardest coding, long-running agents, and judgment-heavy review.
Input
Output
Input
≈ $3.75-$7.5
Output
≈ $18.75-$37.5
Cache
≈ $0.375-$0.75

Anthropic Claude Sonnet 4.6 — a balanced flagship model delivering strong reasoning at production speed and cost, for large-scale, stable deployment
Input
Output
Input
≈ $1.5-$3
Output
≈ $7.5-$15
Cache
≈ $1-$2
APIAny lists public chat and text models on one OpenAI-compatible endpoint, with live credit prices on each model page.
By APIAny Editorial · Updated
Filter by type or provider, compare credit prices, then open a model page for playground tests and request examples.
Sources
The quotations below are from official API documentation that APIAny implements against.
The Chat Completions API endpoint will generate a model response from a list of messages comprising a conversation.
The Gemini API provides access to Google's most capable generative AI models.
The Images API provides several endpoints that let you generate images from text prompts or create edits of existing images.
The APIAny models catalog is the live list of public chat, image, video, audio, and safety models you can call with one API key.