DeepSeek V4.1 Flash
DeepSeek V4.1 Flash: multimodal reasoning, coding and tool use with a 1M-token context window. Separate from DeepSeek V4 Flash.
Input
Output
Input
≈ $0.375-$0.75
Output
≈ $1.5-$3
Cache Read
≈ $0.0075-$0.015
20 Gruppen
DeepSeek V4.1 Flash: multimodal reasoning, coding and tool use with a 1M-token context window. Separate from DeepSeek V4 Flash.
Input
Output
Input
≈ $0.375-$0.75
Output
≈ $1.5-$3
Cache Read
≈ $0.0075-$0.015
DeepSeek V4 Flash — latenzarme Variante von DeepSeek V4, behält die Kern-Reasoning-Fähigkeiten bei und senkt Antwortzeit und Kosten für hohe Nebenläufigkeit.
Input
Output
Input
≈ $0.525-$1.05
Output
≈ $1.57-$3.15
Cache Read
≈ $0.0175-$0.035
GLM 5.3 Flash: multimodal reasoning, coding and tool use with a 1M-token context window. Thinking is always enabled with low, high or max effort.
Input
Output
Input
≈ $0.14-$0.28
Output
≈ $0.49-$0.98
Cache Read
≈ $0.0405-$0.081
Google Gemini 3.5 Flash — schnelles, kosteneffizientes multimodales Chat-Modell mit Bild- und Texteingabe und einer Million Token Kontextfenster.
Input
Output
Input
≈ $0.938-$1.88
Output
≈ $6.25-$12.5
Cache Read
≈ $0.0938-$0.188
MiniMax M3 — großes Sprachmodell mit einem Kontextfenster von einer Million Token, stark im Verständnis langer Dokumente, Reasoning und Tool-Nutzung.
Input
Output
Input
≈ $1.2-$2.4
Output
≈ $5-$10
Cache Read
≈ $1.5-$3
DeepSeek V4 — auf Reasoning fokussiertes Sprachmodell mit 128K-Token-Kontextfenster, stark bei Codegenerierung und komplexer Logik, sehr günstig.
Input
Output
Input
≈ $0.212-$0.425
Output
≈ $0.475-$0.95
Cache Read
≈ $0.025-$0.05
OpenAI GPT-5.5 — Flaggschiffmodell für komplexes Reasoning, Codegenerierung und mehrstufige Anweisungen, mit 400K Token Kontextfenster für anspruchsvolle Apps.
Input
Output
Input
≈ $4.25-$8.5
Output
≈ $25-$50
Cache Read
≈ $0.425-$0.85
OpenAI GPT-5.6 Sol ist die leistungsstarke GPT-5.6-Stufe für komplexes Reasoning, Codegenerierung, Agenten und anspruchsvolle Produktionsautomatisierung.
Input
Output
Input
≈ $6-$12
Output
≈ $40-$80
Cache Read
≈ $0.85-$1.7
OpenAI GPT-6 — a flagship model for complex reasoning, code generation, and multi-step instructions, with a 400K-token context window for demanding apps
Input
Output
Input
≈ $3.5-$7
Output
≈ $30-$60
Cache Read
≈ $0.531-$1.06
xAI Grok 4.5 ist ein leistungsstarkes Modell für Reasoning, Codeanalyse, lange Kontexte und toolgestützte Assistenten.
Input
Output
Input
≈ $0.6-$1.2
Output
≈ $1.2-$2.4
Cache Read
≈ $0.6-$1.2
Grok 4.6 is an advanced AI model designed to deliver fast, writing, coding, research, data analysis, and creative problem-solving through natural conversation
Input
Output
Input
≈ $0.6-$1.2
Output
≈ $1.2-$2.4
Cache Read
≈ $0.6-$1.2
A premium reasoning route for visual front-end prototyping, repository-scale coding, large evidence sets, long-running agents, and complex knowledge work that benefits from a 1.05M-token working context.
Input
Output
Input
≈ $2-$4
Output
≈ $10-$20
Cache Read
≈ $0.2-$0.4
Google Gemini 2.5 Flash Lite — extrem günstiges, latenzarmes Chat-Modell mit einer Million Token Kontextfenster, ideal für hochfrequente Alltagsaufgaben.
Input
Output
Input
≈ $0.1-$0.2
Output
≈ $0.15-$0.3
Cache Read
≈ $0.01-$0.02
Google Gemini 3.1 Flash Lite — verbessertes Reasoning gegenüber der 2.5-Generation bei günstigem Preis, balanciert Geschwindigkeit und Qualität.
Input
Output
Input
≈ $0.15-$0.3
Output
≈ $1-$2
Cache Read
≈ $0.0125-$0.025
Google Gemini 3.1 Pro — multimodales Flaggschiffmodell von Gemini mit starkem komplexem Reasoning und langer Kontextverarbeitung für Produktivanwendungen.
Input
Output
Input
≈ $1.25-$2.5
Output
≈ $7.5-$15
Cache Read
≈ $0.625-$1.25
OpenAI GPT-4o mini — schnelles, günstiges multimodales Chat-Modell mit zügigen Antworten, ideal für alltägliche Fragen und leichte Programmierhilfe.
Input
Output
Input
≈ $0.275-$0.55
Output
≈ $0.412-$0.825
Cache Read
≈ $0.0138-$0.0275
OpenAI GPT-5.4 — leistungsstarkes Modell für fortgeschrittenes Reasoning, Codegenerierung und Agenten-Workflows, mit 400K Token Kontextfenster.
Input
Output
Input
≈ $2.5-$5
Output
≈ $15-$30
Cache Read
≈ $0.25-$0.5
Anthropic Claude Opus 4.8 — Flaggschiffmodell von Claude für anspruchsvollste Reasoning-, Coding- und Langtextaufgaben, mit exzellenter Kontextverarbeitung.
Input
Output
Input
≈ $2.5-$5
Output
≈ $12-$24
Cache Read
≈ $0.25-$0.5
Anthropic Claude Opus 5 from APIAny — Anthropic’s newest Opus-tier flagship for the hardest coding, long-running agents, and judgment-heavy review.
Input
Output
Input
≈ $3.75-$7.5
Output
≈ $18.75-$37.5
Cache Read
≈ $0.375-$0.75
Anthropic Claude Sonnet 4.6 — ausgewogenes Flaggschiffmodell mit starkem Reasoning bei produktionsreifer Geschwindigkeit und Kosten für stabile Großeinsätze.
Input
Output
Input
≈ $1.5-$3
Output
≈ $7.5-$15
Cache Read
≈ $1-$2
APIAny listet öffentliche Chat- und Textmodelle auf einem OpenAI-kompatiblen Endpunkt, mit aktuellen Credit-Preisen auf jeder Modellseite.
Von APIAny Editorial · Aktualisiert
Filtern Sie nach Typ oder Anbieter, vergleichen Sie Credit-Preise und öffnen Sie eine Modellseite für Playground-Tests und Request-Beispiele.
Quellen
Die Zitate stammen aus offizieller API-Dokumentation, gegen die APIAny implementiert.
The Chat Completions API endpoint will generate a model response from a list of messages comprising a conversation.
The Gemini API provides access to Google's most capable generative AI models.
The Images API provides several endpoints that let you generate images from text prompts or create edits of existing images.
Der APIAny-Modellkatalog ist die Live-Liste öffentlicher Chat-, Bild-, Video-, Audio- und Sicherheitsmodelle, die Sie mit einem API-Key aufrufen.