Pay only for what
you use.
No subscriptions. No minimums. Credits never expire. Every component of every call — LLM tokens, TTS characters, STT minutes, telephony time — billed at provider cost with a transparent margin disclosed at the per-call breakdown.
How it works.
One balance. Drawn down per call. Topped up automatically when it crosses your threshold. Nothing else to think about.
Pay-As-You-Go Credits
Purchase credits. Use them at your own pace. Credits never expire. No monthly fee, no seat tax, no minimum commitment.
Auto-Recharge
Set a balance threshold. We top up the configured amount the moment it crosses. You never run out mid-campaign.
Real-Time Deduction
Every call shows its decomposition the moment it ends — LLM, TTS, STT, telephony, line by line. No invoice surprises at month end.
What does a call cost?
Live rates, pulled from each provider every hour. Pick a stack, adjust the duration. The number on the right is what you’ll pay for a single conversation.
Equivalent to $0.079 per minute of conversation.
Language models.
Pricing per 1 million tokens, in / out. Every model that runs on Talkif. Function calling, vision, JSON mode capabilities marked.
| Model | Context | Capabilities | Input | Output |
|---|---|---|---|---|
| Claude Opus 4.8Most intelligent generally available Claude. Best for complex reasoning, agents, and coding tasks. | 1M | FNVISSTREAM | $5.00per 1M tok | $25.00per 1M tok |
| Claude Sonnet 4.5Best balance of speed and intelligence. Fast responses for voice AI. | 200K | FNVISSTREAM | $3.00per 1M tok | $15.00per 1M tok |
| Claude Sonnet 5Latest Sonnet generation. Best combination of speed and intelligence for voice AI. | 1M | FNVISSTREAM | $2.00per 1M tok | $10.00per 1M tok |
| Claude Haiku 4.5Fastest Claude with near-frontier intelligence. Great for high-volume voice AI. | 200K | FNVISSTREAM | $1.00per 1M tok | $5.00per 1M tok |
| Model | Context | Capabilities | Input | Output |
|---|---|---|---|---|
| Gemini 3.5 FlashMost intelligent Gemini built for speed. Frontier intelligence with superior grounding. | 1M | FNVISSTREAM | $1.50per 1M tok | $9.00per 1M tok |
| Gemini 3.1 Pro PreviewMost powerful Gemini model. Best for multimodal understanding and complex agentic tasks. | 1M | FNVISSTREAM | $2.00per 1M tok | $12.00per 1M tok |
| Gemini 3.1 Flash LiteSmall, cost-effective Gemini 3 tier. Ideal for high-volume voice applications. | 1M | FNVISSTREAM | $0.25per 1M tok | $1.50per 1M tok |
| Gemini 3 Pro PreviewPowerful Gemini preview model for multimodal understanding and complex agentic tasks. | 1M | FNVISSTREAM | $2.00per 1M tok | $12.00per 1M tok |
| Gemini 3 Flash PreviewFast and intelligent. Excellent for voice AI with superior search and grounding. | 1M | FNVISSTREAM | $0.50per 1M tok | $3.00per 1M tok |
| Gemini 2.5 ProState-of-the-art multipurpose model. Excels at coding and complex reasoning. | 1M | FNVISSTREAM | $1.25per 1M tok | $10.00per 1M tok |
| Gemini 2.5 FlashHybrid reasoning model with thinking budgets. Great for voice AI with 1M context. | 1M | FNVISSTREAM | $0.30per 1M tok | $2.50per 1M tok |
| Gemini 2.5 Flash LiteSmallest and most cost-effective. Ideal for high-volume voice applications. | 1M | FNVISSTREAM | $0.10per 1M tok | $0.40per 1M tok |
| Model | Context | Capabilities | Input | Output |
|---|---|---|---|---|
| Grok 4.5Most intelligent and fastest Grok. xAI's recommended default for chat and code. | 500K | FNVISSTREAM | $2.00per 1M tok | $6.00per 1M tok |
| Grok 4.3Fast, cost-efficient Grok with 1M context. Strong price-performance for voice AI. | 1M | FNVISSTREAM | $1.25per 1M tok | $2.50per 1M tok |
| Grok 4.1 Fast ReasoningLatest fast model with reasoning. Best price-performance ratio with 2M context. | 2M | FNVISSTREAM | $0.02per 1M tok | $0.05per 1M tok |
| Grok 4.1 FastLatest fast model without reasoning overhead. Ultra-fast responses with 2M context. | 2M | FNVISSTREAM | $0.02per 1M tok | $0.05per 1M tok |
| Grok 4 Fast ReasoningFast model with reasoning capabilities. Great for complex voice AI tasks. | 2M | FNVISSTREAM | $0.02per 1M tok | $0.05per 1M tok |
| Grok 4 FastFast model optimized for speed. Excellent for real-time voice interactions. | 2M | FNVISSTREAM | $0.02per 1M tok | $0.05per 1M tok |
| Grok 4Flagship model with highest intelligence. Best for complex reasoning tasks. | 256K | FNVISSTREAM | $0.30per 1M tok | $1.50per 1M tok |
| Grok 3Previous generation flagship. Strong capabilities at competitive pricing. | 131K | FNSTREAM | $0.30per 1M tok | $1.50per 1M tok |
| Grok 3 MiniLightweight and fast. Good balance of cost and capability for voice AI. | 131K | FNSTREAM | $0.03per 1M tok | $0.05per 1M tok |
| Model | Context | Capabilities | Input | Output |
|---|---|---|---|---|
| GPT-5.6 SolFlagship GPT-5.6 tier. Best for complex reasoning, coding, and creative tasks. | 1.0M | FNVISSTREAM | $5.00per 1M tok | $30.00per 1M tok |
| GPT-5.6 TerraBalanced GPT-5.6 tier. Strong intelligence at mid-tier cost and latency. | 1.0M | FNVISSTREAM | $2.50per 1M tok | $15.00per 1M tok |
| GPT-5.6 LunaFast, cost-efficient GPT-5.6 tier. Great fit for low-latency voice AI. | 1.0M | FNVISSTREAM | $1.00per 1M tok | $6.00per 1M tok |
| GPT-5.5Previous flagship generation with excellent reasoning and instruction following. | 400K | FNVISSTREAM | $5.00per 1M tok | $30.00per 1M tok |
| GPT-5.4Capable general-purpose model with strong reasoning at mid-tier cost. | 400K | FNVISSTREAM | $2.50per 1M tok | $15.00per 1M tok |
| GPT-5.4 MiniFast, affordable small model. Good balance of speed and capability for voice AI. | 400K | FNVISSTREAM | $0.75per 1M tok | $4.50per 1M tok |
| GPT-5.4 NanoCheapest, lowest-latency GPT tier. High-volume, simple conversational tasks. | 400K | FNVISSTREAM | $0.20per 1M tok | $1.25per 1M tok |
| GPT-5.2Highly capable GPT model for complex reasoning, coding, and creative tasks. | 128K | FNVISSTREAM | $1.75per 1M tok | $14.00per 1M tok |
| GPT-5.1High capability model with excellent reasoning and instruction following. | 128K | FNVISSTREAM | $1.25per 1M tok | $10.00per 1M tok |
| GPT-5Powerful model for general-purpose tasks with strong reasoning. | 128K | FNVISSTREAM | $1.25per 1M tok | $10.00per 1M tok |
| GPT-5 MiniFast and cost-effective. Great balance of speed and capability for voice AI. | 128K | FNVISSTREAM | $0.25per 1M tok | $2.00per 1M tok |
| GPT-5.2 Chat LatestAlways points to latest GPT-5.2 chat model. Auto-updates to newest version. | 128K | FNVISSTREAM | $1.75per 1M tok | $14.00per 1M tok |
| GPT-5.1 Chat LatestAlways points to latest GPT-5.1 chat model. Auto-updates to newest version. | 128K | FNVISSTREAM | $1.25per 1M tok | $10.00per 1M tok |
| GPT-5 Chat LatestAlways points to latest GPT-5 chat model. Auto-updates to newest version. | 128K | FNVISSTREAM | $1.25per 1M tok | $10.00per 1M tok |
| GPT-4.1Strong general-purpose model. Good for most conversational AI tasks. | 128K | FNVISSTREAM | $2.00per 1M tok | $8.00per 1M tok |
| GPT-4.1 MiniFast and affordable. Excellent for voice AI with good capability. | 128K | FNVISSTREAM | $0.40per 1M tok | $1.60per 1M tok |
| GPT-4.1 NanoFastest GPT-4 variant. Ideal for simple voice interactions with minimal latency. | 128K | FNSTREAM | $0.10per 1M tok | $0.40per 1M tok |
| GPT-4oVersatile multimodal model. Good for voice AI with vision and function calling. | 128K | FNVISSTREAM | $2.50per 1M tok | $10.00per 1M tok |
| GPT-4o MiniSmall and fast multimodal model. Cost-effective for voice AI workloads. | 128K | FNVISSTREAM | $0.15per 1M tok | $0.60per 1M tok |
Speech to text.
Per minute of audio transcribed. Telephony-tuned models marked. Streaming & diarization where available.
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Amazon TranscribeAmazon Transcribe streaming speech-to-text — real-time recognition with word-level timestamps and telephony (8 kHz) support across 30+ languages. | 23 | STREAMTELMULTI | $0.0240per minute |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Ink WhisperFastest, most affordable streaming STT optimized for real-time voice agents. Handles background noise, telephony artifacts, accents, and domain-specific terminology. | 99 | STREAMNOISETEL | $0.0022per minute |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Nova 3Deepgram's most powerful speech-to-text model. Best accuracy across 70+ languages with real-time streaming. | 74 | STREAM | $0.0092per minute |
| Nova 3 MedicalNova 3 optimized for medical terminology and healthcare conversations. | 08 | STREAM | $0.0077per minute |
| Nova 2High accuracy speech recognition across 40+ languages. Good balance of speed and accuracy. | 48 | STREAM | $0.0058per minute |
| Nova 2 Conversational AIOptimized for conversational AI applications. Low latency for voice assistants and chatbots. | 06 | STREAM | $0.0058per minute |
| Nova 2 Phone CallOptimized for phone call audio. Handles telephony-quality audio with background noise. | 06 | STREAMTEL | $0.0058per minute |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Scribe v2 RealtimeFastest and most accurate live speech recognition. 150ms latency, 90+ languages, automatic language detection. | 74 | STREAMMULTI | $0.0042per minute |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Soniox RT v3Real-time speech-to-text with automatic language identification, speaker diarization, and translation support. | 60 | STREAMDIARMULTI | $0.0020per minute |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Grok STTReal-time speech-to-text with word-level timestamps, speaker diarization, and telephony (µ-law) support. Top accuracy on phone-call benchmarks across 25 languages. | 24 | STREAMDIARNOISETELMULTI | $0.0033per minute |
Text to speech.
Per thousand characters synthesised. Streaming + style + emotion controls noted per model.
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Sonic TurboAll the power of Sonic with half the latency (as low as 40ms). Best for real-time conversational AI. | 15 | STREAM | $0.0374per 1k chars |
| Sonic 2Ultra-realistic speech with accurate transcript following, minimal hallucinations, and excellent voice cloning. | 15 | STREAM | $0.0374per 1k chars |
| Sonic 3Latest streaming TTS with emotion, laughter, speed, volume controls. 42 languages supported. | 42 | STREAM | $0.0374per 1k chars |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Eleven Flash v2.5Our ultra low latency model in 32 languages. Ideal for conversational use cases. | 32 | STREAM | $0.0500per 1k chars |
| Eleven Turbo v2.5Our high quality, low latency model in 32 languages. Best for developer use cases where speed matters. | 32 | STREAM | $0.0500per 1k chars |
| Eleven Multilingual v2Our most life-like, emotionally rich model in 29 languages. Best for voice overs, audiobooks, post-production. | 29 | STREAMSTYLE | $0.1000per 1k chars |
| Eleven Flash v2Our ultra low latency model in English. Ideal for conversational use cases. | 01 | STREAM | $0.1000per 1k chars |
| Eleven Turbo v2Our English-only, low latency model. Best for developer use cases where speed matters and you only need English. | 01 | STREAM | $0.1000per 1k chars |
| Eleven v3 (alpha)The most expressive model. Supports 70+ languages. Requires more prompt engineering than our previous models. | 74 | STREAM | $0.3000per 1k chars |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Gemini 2.5 Flash TTSLow-latency Gemini TTS with natural, promptable voice control. 30 named voices, 23 languages. | 23 | STREAM | $0.0170per 1k chars |
| Gemini 2.5 Pro TTSHighest-quality Gemini TTS for expressive, style-prompted speech. 30 named voices, 23 languages. | 23 | STREAM | $0.0340per 1k chars |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Rime Mist v2Rime's low-latency streaming TTS, purpose-built for real-time voice agents and telephony. WebSocket synthesis emits word-level timestamps. 300+ voices. | 02 | STREAM | $0.0300per 1k chars |
| Rime ArcanaRime's premium, ultra-expressive conversational model with natural prosody and English/Spanish/French/German voices. Higher latency than Mist v2; best when expressiveness matters more than speed. | 04 | STREAM | $0.0400per 1k chars |
| Model | Languages | Capabilities | Rate |
|---|---|---|---|
| Grok TTSExpressive low-latency text-to-speech in 5 voices across 20 languages. Streaming synthesis emits word-level timestamps and supports telephony (µ-law 8 kHz) output. Ideal for real-time voice agents. | 20 | STREAMSPEED | $0.0042per 1k chars |
Telephony.
Per minute carriage. Twilio, SIP/BYOC via Kamailio, WebRTC for in-browser. Phone number rental & recording storage billed separately.
| Provider · route | Direction | Notes | Rate |
|---|---|---|---|
| SIP | inbound | — | $0.0040per minute |
| SIP | outbound | — | $0.0040per minute |
| Twilio | inbound | — | $0.0085per minute |
| Twilio | outbound | — | $0.0140per minute |
| inbound | — | $0.0000per minute | |
| outbound | — | $0.0000per minute |
Phone numbers
From $1.15 per number per month. Monthly cost varies by country and number type. Price shown is for US local numbers.
Call recording
$0.5000 per GB per month. Billed based on actual storage duration. Minimum 1 day charged.
Function execution
Free for the first 10,000 function invocations / month. $0.20 / 10k after. Webhook delivery is always free.
First call is on us.
Sign up — we credit your balance with $5 to test. Enough for roughly 25 qualifying calls across any provider mix.