Text to Speech API for Developers: Piper vs ElevenLabs - NextGenBeing
Back to discoveries

Text to Speech API for Developers: Piper vs ElevenLabs

Text to speech API for developers: a local Piper markdown-to-MP3 pipeline versus ElevenLabs, with measured speed and cost per article.

AI Tutorials 6 min read
Bekzod Erkinov

Bekzod Erkinov

Oct 6, 2026 • 1 views
Text to Speech API for Developers: Piper vs ElevenLabs
Photo by imgix on Unsplash
Size:
Height:
📖 6 min read 📝 4,008 words 👁 Focus mode: ✨ Eye care:

Listen to Article

Loading...
0:00 / 0:00
0:00 0:00
Low High
0% 100%
⏸ Paused ▶️ Now playing... Ready to play ✓ Finished
Table of contents · 7 sections

Text to Speech API for Developers: Piper vs ElevenLabs

Piper renders a markdown article at about 9 to 10x real time on a laptop CPU (RTF 0.10 to 0.11), for free. The same 1,000-word article costs $0.23 to $0.47 on the ElevenLabs API at list price. Below is a pipeline I ran end to end with Piper that cleans markdown, caches audio per chunk and writes one MP3, then the ElevenLabs version of the same loop, the cost math and a blind test to run before you pay.

Some links in this post are affiliate links; see our disclosure.

Set up Piper for markdown to speech

  • Windows 11, 8 logical cores, CPU only, Python 3.12.10 in a venv.
  • pip install piper-tts imageio-ffmpeg, then pip freeze | grep -iE "piper|onnxruntime|imageio":
imageio-ffmpeg==0.6.0
onnxruntime==1.30.0
piper-tts==1.8.0

imageio-ffmpeg bundles its own ffmpeg, so you don't need a system install.

  • Voice: en_US-lessac-medium (a 63 MB .onnx file plus a JSON config):
python -m piper.download_voices --download-dir voices en_US-lessac-medium

Piper lives at OHF-Voice/piper1-gpl. It is GPL-3.0: fine as a build step for your own audio, a licensing conversation if you bundle it in a distributed product.

The test input is a 679-word markdown post (webhook retry budgets) with front matter, an image, a link, two code blocks, a table, a blockquote and lists.

Convert markdown to speech text with Python, and cache TTS audio per chunk

The decisions behind narrate.py:

  • Cleaning: front matter and images are dropped; links and inline code keep their text; table rows become comma-joined sentences; headings get a full stop so the engine pauses. Code blocks become "Code example omitted."
  • Chunking: split on paragraphs, then pack whole sentences up to 400 characters. Chunks never merge across paragraphs, so a chunk's hash changes only when its own paragraph does.
  • Cache key: SHA-256 of the voice filename plus the chunk text, so switching voices can't reuse old audio.
"""narrate.py: markdown -> clean text -> chunks -> cached Piper WAVs -> one MP3."""
import hashlib, re, subprocess, sys, time, wave
from pathlib import Path
import imageio_ffmpeg
from piper import PiperVoice

LIMIT = 400  # max characters per chunk

def clean(md):
    md = re.sub(r"\A---.*?---\s*", "", md, flags=re.S)             # front matter
    md = re.sub(r"```.*?```", " Code example omitted. ", md, flags=re.S)
    md = re.sub(r"!\[[^\]]*\]\([^)]*\)", "", md)                    # images
    md = re.sub(r"\[([^\]]+)\]\([^)]*\)", r"\1", md)                # links -> text
    md = re.sub(r"`([^`]+)`", r"\1", md)                            # inline code
    md = re.sub(r"^\|[-| ]+\|$", "", md, flags=re.M)                # table rules
    md = re.sub(r"^\|(.*)\|$",                                      # table rows
                lambda m: ", ".join(c.strip() for c in m.group(1).split("|")) + ".",
                md, flags=re.M)
    md = re.sub(r"^#{1,6}\s*(.+)$", r"\1.", md, flags=re.M)         # headings -> sentences
    md = re.sub(r"^>\s?", "", md, flags=re.M)                       # blockquotes
    md = re.sub(r"^\s*(?:[-*]|\d+\.)\s+", "", md, flags=re.M)       # list markers
    md = re.sub(r"(\*{1,3})([^*\n]+?)\1", r"\2", md)                # *italic*, **bold**
    md = re.sub(r"(?<!\w)(_{1,3})([^_\n]+?)\1(?!\w)", r"\2", md)    # _italic_, not snake_case
    return re.sub(r"\.\.+", ".", md)

def paragraphs(text):
    return [re.sub(r"\s+", " ", p).strip()
            for p in re.split(r"\n\s*\n", text) if p.strip()]

def chunk(paras, limit=LIMIT):
    out = []
    for p in paras:                      # never merge across paragraphs
        cur = ""
        for s in re.split(r"(?<=[.!?])\s+", p):
            if cur and len(cur) + 1 + len(s) > limit:
                out.append(cur); cur = s
            else:
                cur = f"{cur} {s}".strip()
        if cur:
            out.append(cur)
    return out

def main(md_path, voice_path, cache_dir, out_mp3):
    t0 = time.perf_counter()
    cache = Path(cache_dir); cache.mkdir(exist_ok=True)
    voice = PiperVoice.load(voice_path)
    t_load = time.perf_counter() - t0

    chunks = chunk(paragraphs(clean(Path(md_path).read_text(encoding="utf-8"))))
    paths, rendered, render_s = [], 0, 0.0
    for c in chunks:
        key = hashlib.sha256(f"{Path(voice_path).name}|{c}".encode()).hexdigest()[:16]
        p = cache / f"{key}.wav"
        if not p.exists():
            t = time.perf_counter()
            with wave.open(str(p), "wb") as w:
                voice.synthesize_wav(c, w)
            render_s += time.perf_counter() - t
            rendered += 1
        paths.append(p)

    t = time.perf_counter()              # join with 250 ms of silence
    with wave.open(str(paths[0])) as w0:
        params = w0.getparams()
    gap = b"\x00" * (int(params.framerate * 0.25) * params.sampwidth * params.nchannels)
    wav_path = Path(out_mp3).with_suffix(".wav")
    with wave.open(str(wav_path), "wb") as out:
        out.setparams(params)
        for p in paths:
            with wave.open(str(p)) as w:
                assert w.getparams()[:3] == params[:3], "mixed voices in cache"
                out.writeframes(w.readframes(w.getnframes()) + gap)
    t_join = time.perf_counter() - t

    t = time.perf_counter()              # 64 kbps mono MP3
    subprocess.run([imageio_ffmpeg.get_ffmpeg_exe(), "-y", "-loglevel", "error",
                    "-i", str(wav_path), "-b:a", "64k", out_mp3], check=True)
    t_mp3 = time.perf_counter() - t

    print(f"chunks={len(chunks)} rendered={rendered} cached={len(chunks)-rendered} "
          f"chars={sum(map(len, chunks))}\n"
          f"load={t_load:.1f}s render={render_s:.1f}s join={t_join:.1f}s mp3={t_mp3:.1f}s "
          f"total={time.perf_counter()-t0:.1f}s")

if __name__ == "__main__":
    main(*sys.argv[1:5])

Run it with python narrate.py article.md voices/en_US-lessac-medium.onnx cache out.mp3.

What clean() does to the input

Raw markdown from the sample (a table and a line with inline code, a link and an image):

| Failure | Retry? | Why |
|---|---|---|
| `429` | Yes, honor `Retry-After` | The sender is telling you the pace |
| `500`, `503` | Yes, with budget | Probably transient |

Webhook consumers fail. A `503` for ninety seconds. See [the SRE book](https://x.y/z).
![d](a.png)

What the engine receives (real output of clean() on that table, and on the second snippet):

Failure, Retry?, Why.

429, Yes, honor Retry-After, The sender is telling you the pace.
500, 503, Yes, with budget, Probably transient.

Webhook consumers fail. A 503 for ninety seconds. See the SRE book.

The pipes, backticks, URL and image are gone, and each table row is now one sentence.

One pitfall I hit in review: my first emphasis regex, [*_]{1,3}([^*_]+)[*_]{1,3}, also matches the underscores in snake_case. Because inline-code backticks are stripped first, max_retry_count became maxretrycount. The script above only strips underscores at word edges. Real output of clean() on Set `max_retry_count` to 5, **never** zero, and *please* _log_ it.:

Set max_retry_count to 5, never zero, and please log it.

Piper may still read the underscores oddly, so add identifiers like max_retry_count to the pronunciation table below if your docs are full of them.

On the sample, 679 words of markdown became 574 spoken words and 3,355 characters, which is 5.84 characters per word. I use that ratio for pricing below.

The numbers on this machine

Three runs of the script above, same article, fresh cache directory. Run 3 has one sentence edited. Real output of run 1:

chunks=25 rendered=24 cached=1 chars=3355
load=2.1s render=22.9s join=0.0s mp3=1.3s total=26.3s
Run Chunks rendered Cached Model load Render WAV join MP3 encode Total
1. Cold cache 24 1 2.1 s 22.9 s 0.0 s 1.3 s 26.3 s
2. No edits 0 25 2.5 s 0 s 0.0 s 0.9 s 3.5 s
3. One sentence edited 1 24 2.3 s 2.2 s 0.0 s 1.2 s 5.8 s

The fixed overhead is about 3 seconds: loading the 63 MB model plus the MP3 encode. Per-chunk cost is what the cache removes: the cold run spends 22.9 s rendering, the edited run 2.2 s. The first model load after a long idle period can be much slower: on a 4-chunk test file it took 8.7 to 9.8 s.

Two cold runs of this article gave a real-time factor (render time over audio length) of 0.10 to 0.11, about 9 to 10x faster than real time, at 146 to 167 characters of text per second (run 1: 22.9 s of rendering for 204.7 s of audio, RTF 0.112, 146 chars/s; the earlier run: 0.098, 167 chars/s). One chunk in run 1 was already a cache hit: two chunks had identical text and therefore the same hash.

Sizes: the concatenated WAV is about 9.0 MB (22,050 Hz, 16-bit, mono); the 64 kbps mono MP3 is about 1.6 MB for 3 minutes 25 seconds of audio. To check your own output, ffmpeg prints the duration (imageio-ffmpeg ships no ffprobe, so run the bundled ffmpeg with -i and no output file). Real output on a small 5-chunk sample file, not the full article:

python -c "import imageio_ffmpeg; print(imageio_ffmpeg.get_ffmpeg_exe())"   # path to ffmpeg
ffmpeg -i out.mp3 2>&1 | grep Duration
ls -l out.mp3
  Duration: 00:00:18.60, start: 0.050113, bitrate: 64 kb/s
-rw-r--r-- 1 Bekzod 197121 149045 Oct  7 03:06 out.mp3

At 64 kbps that's 8 KB per second, so 3:25 of audio works out to about 1.6 MB.

A quality proxy: how Piper reads acronyms

I didn't listen to the output, but I can show what Piper's front end will say. voice.phonemize(word) returns the phonemes it will speak. For tokens from the sample article:

Text Phonemes (IPA) Spoken as
SRE ˌɛsˌɑːɹɹˈiː "S R E", letter by letter
API ˌeɪpˌiːʲˈaɪ "A P I"
HTTP ˌeɪtʃtˌiːtˌiːpˈiː "H T T P"
AWS ˈɔːz "awz", read as a word
IDs aɪdˈiː "eye-dee", plural dropped
ONNX ˈɑːŋŋks a guess
ElevenLabs ᵻlˈɛvən lˈæbz "eleven labs"
503 fˈaɪvhˈʌndɹɪd θɹˈiː "five hundred three"

Common acronyms are spelled out, but AWS is mangled and IDs loses its plural. That's two misses among the eight tokens I checked, judged from phonemes, not by ear. I haven't tested any API model on these, so I can't say whether it does better. A pronunciation table is a cheap fix that works for either engine. Run it after clean(); the \b word boundaries keep AWS from matching inside AWSome:

PRONOUNCE = {"AWS": "A W S", "IDs": "I Ds", "ONNX": "onyx"}

def pronounce(text):
    for word, spoken in PRONOUNCE.items():
        text = re.sub(rf"\b{re.escape(word)}\b", spoken, text)
    return text

Real output of pronounce(clean("Use AWS request IDs with ONNX. AWSome stays.")), re-run with the boundaries in place:

Use A W S request I Ds with onyx. AWSome stays.

Piper's sample page lets you audition voices before downloading one.

Cost and limits of the ElevenLabs text to speech API

The ElevenLabs code below was never run against the live API; I had no key. I did run it against a stub requests.post that returns a real MP3 tone and fake request-id headers. That tests the caching, chaining and ffmpeg join logic, not the real responses or the audio quality.

From the convert reference:

  • POST https://api.elevenlabs.io/v1/text-to-speech/{voice_id} with the xi-api-key header.
  • Body: text (required), model_id (default eleven_multilingual_v2), optional voice_settings, previous_text, next_text, previous_request_ids, next_request_ids.
  • output_format is a query parameter, not a body field. Default mp3_44100_128.
  • At most 3 request IDs per ID list.

The models page (checked 2026-10-07) lists these per-request character limits:

Model ID Limit Notes
eleven_v4 10,000 90+ languages
eleven_v3 5,000 70+ languages
eleven_multilingual_v2 10,000 29 languages
eleven_flash_v2_5 40,000 32 languages, lower price per character
eleven_flash_v2 30,000 English only

A 1,000-word article (about 5,850 spoken characters) fits in one request on everything except eleven_v3, so most posts need no chunking. Longer ones do, and chunk seams can change prosody. Request stitching addresses that: pass the request-id response header of earlier calls as previous_request_ids. Per the cookbook, the IDs "should be no older than two hours" and stitching "is not available for the eleven_v3 model".

Stitching and a per-chunk cache conflict: a cache hit produces no fresh request ID, and a stitched chunk's audio depends on its neighbour. The loop below handles both. The cache key includes a hash of the previous chunk, so editing paragraph 3 re-renders paragraph 4 too. A chunk that follows a cache hit has no ID to chain from, so it sends previous_text instead. Chunks are merged to a larger size first with group(); on this sample it turns 25 chunks into 3 at a 1,500-character limit.

def group(chunks, limit):
    """Join consecutive chunks with a blank line while the group stays <= limit chars."""
    out = []
    for c in chunks:
        if out and len(out[-1]) + 2 + len(c) <= limit:
            out[-1] += "\n\n" + c
        else:
            out.append(c)
    return out

# Piper path:  chunk(paras)                      -> up to 400 chars each
# API path:    group(chunk(paras), limit=1500)   -> a few paragraphs each
import hashlib, subprocess, time
from pathlib import Path
import requests
import imageio_ffmpeg

def eleven_key(voice_id, model_id, text, prev_text):
    prev = hashlib.sha256(prev_text.encode()).hexdigest()[:8]
    return hashlib.sha256(f"{voice_id}|{model_id}|{prev}|{text}".encode()).hexdigest()[:16]

def render(chunks, api_key, voice_id, cache, out_mp3, model_id="eleven_flash_v2_5"):
    url = f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}"
    cache = Path(cache); cache.mkdir(exist_ok=True)
    files, prev_ids, calls = [], [], 0
    for i, text in enumerate(chunks):
        prev_text = chunks[i - 1] if i else ""
        p = cache / f"{eleven_key(voice_id, model_id, text, prev_text)}.mp3"
        if not p.exists():
            body = {"text": text, "model_id": model_id}
            if prev_ids:                       # previous chunk rendered just now: chain by ID
                body["previous_request_ids"] = prev_ids[-3:]    # max 3
            else:                              # first chunk, or previous one came from cache
                if prev_text: body["previous_text"] = prev_text
            if i + 1 < len(chunks): body["next_text"] = chunks[i + 1]
            for attempt in range(4):
                r = requests.post(url, params={"output_format": "mp3_44100_128"},
                                  headers={"xi-api-key": api_key}, json=body, timeout=120)
                if r.status_code != 429 and r.status_code < 500:
                    break
                time.sleep(2 ** attempt)
            r.raise_for_status()
            p.write_bytes(r.content); calls += 1
            rid = r.headers.get("request-id")
            prev_ids = prev_ids + [rid] if rid else []
        else:
            prev_ids = []                      # cache hit: no fresh ID, chain restarts
        files.append(p)
    lst = cache / "list.txt"
    lst.write_text("".join(f"file '{f.resolve().as_posix()}'\n" for f in files))
    subprocess.run([imageio_ffmpeg.get_ffmpeg_exe(), "-y", "-loglevel", "error", "-f", "concat",
                    "-safe", "0", "-i", str(lst), "-b:a", "64k", out_mp3], check=True)
    return calls

The join decodes each MP3 with ffmpeg's concat demuxer and re-encodes, rather than concatenating raw MP3 bytes, which can leave glitchy seams and wrong duration headers. Get a key at ElevenLabs.

Stub run: four chunks, then an edit to chunk 2. The first line is print(eleven_key("v","m","t","a"), eleven_key("v","m","t","b"), eleven_key("v","m","t","a")==eleven_key("v","m","t","b")); the rest are the logged request bodies minus text and model_id:

869aff910f52ed29 bd56c814645427ec False
cold calls: 4
{'next_text': 'Para two.'}
{'previous_request_ids': ['rid1'], 'next_text': 'Para three.'}
{'previous_request_ids': ['rid1', 'rid2'], 'next_text': 'Para four.'}
{'previous_request_ids': ['rid1', 'rid2', 'rid3']}
edit para 2 calls: 2
{'previous_text': 'Para one.', 'next_text': 'Para three.'}
{'previous_request_ids': ['rid1'], 'next_text': 'Para four.'}

The two keys differ only in the previous chunk, so the equality is False. After the edit, chunk 2 follows a cache hit and uses previous_text; chunk 3's key changed, so it re-renders and chains from chunk 2's fresh ID; chunk 4 is a cache hit. Stitched calls are sequential, and I have no measured wall-clock for a real API render, so time one before relying on a batch job.

ElevenLabs text to speech API cost per article

From the API pricing page, read 2026-10-07 (prices change; recheck before you budget), per 1,000 characters:

Model $/1K chars Per 1,000-word article (~5,850 chars)
Flash / Turbo $0.04 $0.23
v2 Multilingual $0.08 $0.47
v3 $0.08 $0.47
v4, list price $0.08 $0.47
v4, promo, ends Oct 12 $0.022 $0.13
v4 Turbo, list / promo (ends Oct 12) $0.04 / $0.011 $0.23 / $0.06
Piper, local $0 $0 (about 40 s of CPU at my measured rate)

The promo prices end five days after I checked, so budget on list price. The 5,850 figure is the 5.84 characters per word I measured above; your prose will move it. The same page lists a Starter plan at $1 for the first month, then $6 per month, with 272.727k characters of v4 TTS included. Billing is per character, so chunking and stitching cost the same as one request.

Re-rendering is where budgets leak: ten whole-article re-renders for typos at Flash pricing is $2.34, not $0.23. The hash cache pays for itself even on a paid engine.

Piper or ElevenLabs: how to decide

I measured speed, cost and acronym handling. I did not measure how either engine sounds, and quality is the main reason to pay, so run this ten-minute blind test first:

  1. Pick two paragraphs from your docs, one prose and one dense with product names and acronyms.
  2. Render each with Piper and with one ElevenLabs model, using the same text.
  3. Rename the four files in random order, note the mapping, and have two colleagues rank them blind.
  4. If the ranking doesn't track the engine, take the free one.

Then:

  • Piper: drafts, internal docs, changelogs, anything regenerated often, or content that can't leave your network, if GPL-3.0 and 9 to 10x real time on CPU suit you.
  • ElevenLabs: public-facing audio where the blind test showed a preference, or a language Piper has no voice for (audition Piper's voices on its sample page). At 50 articles a month on Flash, that's about $12.

What to do next

  1. Create a venv, install piper-tts imageio-ffmpeg, and download en_US-lessac-medium.
  2. Run clean() on your own markdown and read the output before synthesizing; check acronyms with voice.phonemize().
  3. Keep the hash cache from day one, and include voice, model and previous-chunk hash in the key if you move to an API.
  4. Time one cold render and compute your own characters per second before trusting mine.
  5. Run the blind test before committing to a paid plan.
Bekzod Erkinov

Bekzod Erkinov

Author

Founder of NextGenBeing. Software engineer working with Laravel, Python, and cloud infrastructure. Writes about patterns that actually hold up in production. Based in Tashkent, Uzbekistan.

🎁 Free guide

Get the AI-Assisted Developer's Field Guide

The workflow, prompts, and tools I use to ship faster with AI — free when you subscribe. Plus new deep-dives in your inbox. No spam, unsubscribe anytime.

Comments (0)

Please log in to leave a comment.

Log In

Related Articles

Don't miss the next deep dive

Get one well-researched tutorial in your inbox each week. No spam, unsubscribe anytime.