Akapulu Labs logo Akapulu Labs Blog

Real-Time Talking Avatar Providers: Every Price, Sourced and Dated (September 2026)

Every real-time talking avatar provider in September 2026, with its published price, its latency where one is published, and a blank where it is not. Bring-your-own-stack renderers run $0.01 to $0.18 a minute; managed conversation stacks run $0.11 to $0.37. Every figure was read off the vendor's own live page on 14 September 2026.

Real-Time Talking Avatar Providers: Every Price, Sourced and Dated (September 2026)

If you are choosing a real-time talking avatar provider, the question that decides your shortlist is not "which avatar looks best." It is whether you already have a voice agent. Answer that first and the field splits cleanly in two, with a 10x to 35x price difference between the halves. Most comparison articles never make the split, so Synthesia and a $0.01/min render API end up in the same table.

Below are the real-time talking avatar providers, each with its published price, its published latency where there is one, and a blank where there is not. The short answer first, then the index.

If you already run STT, an LLM and TTS and you only need a synced talking face, you want a bring-your-own-stack renderer. Look at Simli, Beyond Presence Speech-to-Video, or LiveAvatar LITE. Published rates land between roughly $0.01 and $0.18 a minute.

If you have none of that and want one vendor to run the whole conversation, you want a managed stack. Look at Tavus, Anam, D-ID Agents, or Akapulu Labs. Published rates land between roughly $0.11 and $0.37 a minute.

We build one of these, and it is in the table below, scored on the same axes as everyone else including the rows where it loses.

Most avatar comparisons are comparing two different products

There are two categories wearing the same name.

Rendered video. You write a script, you get an MP4 back. Nobody talks to it. This is Synthesia, HeyGen's original API, D-ID Studio. Synthesia's own pricing page sells "minutes of video per month," which is the tell.

Real-time conversational. It joins a live call, listens, and answers. This is Tavus, Anam, Simli, Beyond Presence, LiveAvatar, Akapulu Labs, and until this year Soul Machines.

The two categories do not compete with each other. They do not share a use case, a price structure, or a latency budget, and no buyer is choosing between them. Yet Synthesia and D-ID Studio appear in "best real-time talking avatar" listicles constantly, which is how a shortlist ends up with two products on it that cannot do the same job.

One makes files. The other makes conversations. A comparison that does not say which is which is not comparing anything.

The split inside real-time that actually sets the price

Within real-time there is a second division, and it is the one that moves the number on your invoice by an order of magnitude.

Bring-your-own-stack. You already have speech-to-text, a language model, and text-to-speech running. You send audio, the vendor returns lip-synced video. They do one job: render a face.

Managed stack. The vendor runs the transcription, the model, the voice, and the render. You send a configuration and a scenario, they hand back a call.

One vendor publishes both modes and prices them side by side. Beyond Presence sells Speech-to-Video, which is bring-your-own, and Conversational, which is managed. Their published overage rates:

Beyond Presence tierSpeech-to-Video (you bring the stack)Conversational (they run it)
Starter€0.175/min€0.35/min
Growth€0.10/min€0.20/min
Scale€0.0875/min€0.175/min

Exactly double, at every tier. Same company, same GPUs, same avatars. What separates the two prices is four stages the managed side runs and the renderer does not: transcription, the language model, speech synthesis, and the turn-taking logic that decides when the avatar stops talking because someone interrupted it. Each of those is a vendor with its own margin, and the managed rate carries all four.

So the second half of the bill does not buy better video. It buys not having to assemble the other four stages yourself.

Every real-time talking avatar provider, priced

Every figure below came off the vendor's own live page on 14 September 2026. Where a vendor does not publish a number, the cell says so. No gap here is filled with an estimate or with a figure from another blog post, because nearly all of those are stale. Tavus roundups still circulating quote "$22 for 60 minutes" against a live page that says $59 for 100.

ProviderTypeManaged rateBYO rateFree tierMax concurrencyPublished latency
TavusReal-time, managed$0.37/min → $0.32/min overagen/a25 min3 → 10Not on pricing page
AnamReal-time, managed$0.16 → $0.14 → $0.12 → $0.11/minn/a30 min/mo1 → 10180 ms (vendor-stated)
Beyond PresenceBoth€0.35 → €0.175/min€0.175 → €0.0875/min; €0.03 at scale40 min1 → 50Not on pricing page
LiveAvatar (HeyGen)Real-time, FULL + LITEFULL: 1 credit / 30 sLITE: 1 credit / min; $0.01/min at scale10 credits/moUnlimited on paid tiersNot published
SimliReal-time, BYOn/aPer-minute rate not published. Plans are $0 / $10 / $49 / $249 a month, visible only after signup$10 on signup + 50 min/moNot publishedUnder 300 ms
Akapulu LabsReal-time, managed$0.195 → $0.175/min (derived)None. Akapulu Labs has no BYO mode.10 credits total, not monthly2 → 15150 ms first frame, 290 ms end to end
HedraBothNot publishedn/an/aNot publishedNot published
D-IDBothNot readable on the public pricing pagen/an/aNot publishedNot published
SynthesiaRendered onlyNot a conversational product. Sold as minutes of generated video. Included here so you can stop seeing it in real-time comparisons.
Soul MachinesReal-timeIn receivership since 5 February 2026. Still listed as a live option in most roundups.

Derived figures are marked. Ours is plan price divided by included credits at one credit per live minute: the arithmetic, rather than a rate we do not print on the pricing page.

Converting the credit systems

Three vendors price in credits, which makes direct comparison harder. The arithmetic:

  • LiveAvatar bills FULL mode at 1 credit per 30 seconds and LITE at 1 credit per minute. At $19 for 200 credits that is about $0.19/min FULL, $0.095/min LITE. At $475 for 6,000 credits it is about $0.158 and $0.079.
  • Akapulu Labs bills one credit per live minute: $48.97 for 251 credits is $0.195/min, and $1,049.98 for 6,000 is $0.175/min.
  • Beyond Presence publishes per-minute overages directly, which is the friendliest disclosure in the category.

No, avatars do not cost a cent a minute

You will read that the price of a talking face has collapsed to around $0.01 a minute. The floor has. The list price has not, and the two get quoted interchangeably.

The sub-cent numbers are real but they are enterprise, at-scale, bring-your-own-stack rates: LiveAvatar's "$0.01/min at scale," Beyond Presence's "€0.03/min at scale," both on Enterprise plans with negotiated commitments. What you can actually sign up for this afternoon and put on a card is $0.08 to $0.37 a minute, depending on which half of the split you are in.

The floor is a cent. The list is ten to thirty-five times that. If a vendor quotes you the floor, ask what commitment it is attached to.

What nobody in this category publishes

The gaps are as informative as the numbers, and we are in several of them.

  • Simli has no public pricing page. simli.com/pricing returns a 404 and every pricing link on the site points into the signed-in app. The plans are $0, $10, $49 and $249 a month, which you can only find out by creating an account.
  • LiveAvatar publishes no latency figure, only the phrase "low latency."
  • Nobody publishes self-hosting terms. Not one vendor in this table, Akapulu Labs included. Ours exists on Enterprise and is unpriced, which is the same non-answer everyone else gives.
  • Anam's monthly prices are a JavaScript digit animation, so they are invisible to any tool that does not execute scripts. Their per-minute rates are readable; the monthly figures are not.
  • We do not publish how many avatars are in the catalogue. It sits behind a login.

Where each one is the right answer

Nobody in this table is the right answer for everyone, Akapulu Labs included.

Anam is the best-priced managed stack that is openly documented, at $0.11 to $0.16 a minute with the widest integration surface in the category: ElevenLabs, LiveKit, VideoSDK, Pipecat, Agora, plus website-builder embeds. If you want a managed avatar and you are price-sensitive, start here.

Simli is the cheapest way to put a face on a voice agent you have already built, and it ships templates for OpenAI, ElevenLabs, Vapi and LiveKit. The trade is that you have to ask them what it costs.

Beyond Presence has the clearest pricing page in the category, both product modes side by side, an n8n node, and the highest published concurrency outside LiveAvatar. If you want to compare managed against BYO with real numbers, they are the only vendor who lets you.

Tavus has the widest distribution in the category: the first LiveKit plugin, a maintained Pipecat service, a Vapi provider, real example repos. If you want the path of least integration resistance, it is here. But it is also the most expensive managed rate in the table at $0.32 to $0.37, and it charges a 30-second minimum on every conversation.

LiveAvatar is the concurrency answer. Unlimited concurrent sessions on every paid tier, at no per-session cost, is unmatched here. If you are running many simultaneous conversations, the rest of this table is academic.

D-ID has the enterprise footprint and the named case studies, which is what procurement asks for and what most of this table cannot supply. But the per-minute economics are not readable from the public pricing page.

Where Akapulu Labs fits, and where it loses

Akapulu Labs is a managed stack. Speech-to-text, the model, the voice and the face render run as one service, with the voice synthesis and the rendering in the same container on the same GPU rather than hopping between vendors. Because the stages share a device, their costs hide behind each other instead of summing. Time to first frame: 150 ms. End to end: 290 ms, sustained at 25 frames per second. Both were measured and written up when the model shipped.

The limits, as of September 2026:

  • We have no bring-your-own-stack mode. There is no documented audio-in, video-out endpoint, so if you already run a Pipecat or LiveKit voice agent you cannot bolt us onto it. Use Simli or Beyond Presence Speech-to-Video instead. This is the biggest gap in the product and the first thing we intend to fix.
  • No ecosystem plugins. No Pipecat service, no LiveKit plugin, no ElevenLabs guide, no n8n node. Every other real-time vendor in the table has at least one.
  • The free tier is ten credits, and they are lifetime rather than monthly. Tavus gives 25 minutes, Anam 30 a month, Beyond Presence 40, LiveAvatar 10 credits every month. Ours is a look, not a trial.
  • The price sits mid-to-high. Around $0.175 to $0.195 a minute puts us above Anam and well above anything bring-your-own.
  • Concurrency caps at 15. LiveAvatar says unlimited and Beyond Presence lists 50.
  • We are early. No third-party reviews, no published compliance posture, and model choice is currently OpenAI only.

Better read here than discovered in week two.

Every figure was read off the vendor's own live page on 14 September 2026. Akapulu Labs re-checks every row and re-dates this page each quarter.

Common questions

What are the best real-time talking avatar providers?

There is no single best one, because the category splits in two. If you already run speech-to-text, a language model and text-to-speech, the providers built for you are Simli, Beyond Presence Speech-to-Video and LiveAvatar LITE, at roughly $0.01 to $0.18 a minute. If you want one vendor to run the whole conversation, they are Tavus, Anam, D-ID Agents and Akapulu Labs, at roughly $0.11 to $0.37 a minute. Every price here was read off the vendor's own page on 14 September 2026.

What is the cheapest real-time avatar API?

For bring-your-own-stack, published rates go as low as €0.03 a minute at scale on Beyond Presence and $0.01 a minute at scale on LiveAvatar Enterprise; both are negotiated commitments rather than self-serve. Among managed stacks with public rates, Anam is the lowest at $0.11 to $0.16 a minute.

Do I need a managed stack or just a renderer?

If you already run speech-to-text, a language model and text-to-speech, you need a renderer and you should pay renderer prices. If you do not, a managed stack will cost roughly double and save you the integration.

Is Synthesia a real-time avatar provider?

No. Synthesia generates rendered video from a script. It is sold in minutes of video produced, not minutes of conversation, and it cannot hold one.

Is Soul Machines still available?

Soul Machines entered receivership on 5 February 2026. It still appears in most comparison articles.