Proof of concept for Learn WordPress issue #3649, built by @softglaze. A demonstration to show read-aloud in context. This is not the live Learn WordPress site.
Learn WordPress Search
Learn › Tutorials › Intro to Block Patterns

Updated 10 September, from the feedback on the issue

Tutorial · has video

Intro to Block Patterns

Erica Varlese14 minutesLevel: Intermediate
🔊 Listen to this lesson Free · on your device

The proposed feature: a button that reads the lesson text aloud with the browser's own voice. Speed presets, a volume button and a position slider that seeks to the word. Voice choice is in the menu at the top left.

0:00 0:00
Speed

If you're already familiar with the Block Editor, or have already seen the previous video, today we're going to take a look at Block Patterns, what they are, and of course, how to use them. Let's get started.

A block pattern is a ready-made arrangement of blocks that you can insert into a post or page in one step, then adjust to your needs. Patterns save you from building common layouts, like a call to action, a gallery, or a pricing table, block by block every time.

You add a pattern the same way you add any block, through the inserter, then edit the text and images inside it. Because a pattern is just blocks, everything you already know about moving and styling blocks still applies.


Lesson · text based, no video

Introduction to WordPress

Beginner WordPress UserModule 1 of 7Level: Beginner

Many Learn lessons are written text with no video lecture. Read-aloud matters most here, because there is no narration at all today. The same free button works on plain lesson text.

🔊 Listen to this lesson Free · on your device

A text-only lesson read aloud by the browser. No video was ever recorded for it, so this is the case read-aloud helps most.

0:00 0:00
Speed

Hi, and welcome to this introduction to WordPress. In this lesson, we'll delve into the basics, helping you understand what you can achieve with WordPress, even if you're a beginner.

At its core, WordPress empowers you to create your own website or blog. Millions of websites have been built using WordPress, ranging from online stores and small businesses to renowned entities like NASA, the Walt Disney Company, and the Facebook Newsroom.

WordPress is a CMS, or content management system. In simple terms, it is a free tool that helps you create, edit, and manage your own website or blog without needing to learn code.

About the "browser voice" (the free, native option)

The free option above does not use any online service. It uses the Web Speech API, a text-to-speech engine built into every modern browser (Chrome, Edge, Safari, Firefox). When you press play, the browser reads the text using the voices already installed on your own computer or phone.

Why it suits WordPress.org

What to keep in mind

How the position slider works, since the API has no seek

Worth stating plainly, because it shapes the decision. The Web Speech API is fire and forget: it has no currentTime, no duration and no seek. You hand it text, it speaks, and that is the whole interface. There is nothing to drag a playhead along.

So the slider above is built rather than borrowed. The lesson is split into sentences and each is measured from its word count and the chosen speed. Dragging the slider works out which sentence that time falls in, then starts speaking from the nearest word inside it rather than from the top of the sentence, so nudging forward five seconds moves you five seconds. Progress is then tracked with the API's word boundary events, so the bar follows the actual speech rather than a stopwatch.

Two honest limits remain. The total length is an estimate until the lesson has been spoken once, because the API will not tell you how long anything takes. And every seek restarts an utterance, so there is a small gap of a few dozen milliseconds. A pre-generated audio file has neither problem: it is an ordinary <audio> element with a real duration that seeks to the exact second. One more small argument for the file route on flagship lessons.

Try it with your own device, in any language you have

This page reads the list of voices installed in your browser, so you can test it in whatever languages your device supports. Open the menu on the player below, pick a language and a voice, edit the text if you like, and press play.

Checking which voices are installed in your browser ...

Language and voice are in the menu on the left, which lists the best four voices your device has for the language you pick. Edit the text below if you like.

0:00 0:00
Speed

Proof of concept: how should read-aloud sound?

The lessons above show the free, browser-native feature working. This section compares that free voice against a commercial voice and the lesson's own narrator. For the commercial voice you can switch between several male and female options. Everything here is for comparison only.

English, the same passage

If you're already familiar with the Block Editor, or have already seen the previous video, today we're going to take a look at Block Patterns, what they are, and of course, how to use them. Let's get started.

Browser voice Free · on device

Web Speech API. No service, no account, no cost. Uses the voices already on this device.

0:00 0:00
Speed

Commercial API voice For comparison only

Pre-generated with ElevenLabs. Voice is in the menu on the left, currently Bella.

Original video audio Reference

The lesson's own narrator, Erica Varlese, at 0:33 to 0:44 of the video above.

Non-English: Urdu

Urdu is right-to-left, and most devices ship no Urdu voice at all, so the free browser route below usually reports that it cannot speak it. This is not hypothetical for Learn: the lesson Introduction to WordPress is already translated into Urdu in Learn #3613. The passage below is that reviewed community translation.

سلام، اور WordPress کے اس تعارفی سبق میں خوش آمدید۔ اس سبق میں ہم بنیادی باتوں کا جائزہ لیں گے تاکہ آپ سمجھ سکیں کہ WordPress سے کیا کچھ حاصل کیا جا سکتا ہے، چاہے آپ بالکل نئے ہی کیوں نہ ہوں۔

Browser voice Free · on device

0:00 0:00
Speed

Commercial API voice For comparison only

Pre-generated with ElevenLabs (model Eleven v3, which covers Urdu). Voice is in the menu on the left, currently Bella.

How each sample was produced

Stated so anyone can reproduce or challenge every clip.

SampleRouteService / voicesSettingsCost
Browser voiceWeb Speech APIWhatever is installed on the deviceDefault pitch; adjustable speed and volumeFree
English, commercialElevenLabs, static filesBella, Sarah (female); Eric, George (male)eleven_multilingual_v2; stability 0.5, similarity 0.75; MP3 128 kbpsFree tier
English, referenceFrom the lesson videoErica Varlese (Intro to Block Patterns)0:33.8 to 0:44.3, unedited, mono MP3n/a
Urdu, commercialElevenLabs, static filesBella, Sarah (female); Eric, George (male)eleven_v3; MP3 128 kbpsFree tier

How to make the voice less robotic

This was asked directly on the issue, so here is the full answer rather than a one-liner. The headline: most of the robotic quality is not the browser, it is which voice the browser picked, and by default it usually picks one of the worst ones installed. Fixing that costs nothing.

1. Why it sounds robotic in the first place

Three generations of text-to-speech are all still shipping on real devices, and a browser will happily hand you any of them:

GenerationHow it worksTypical examples still installedHow it sounds
FormantGenerates sound from rules and filters. No human recordings involved at all.eSpeak, eSpeak NG, most Linux defaults, some Android fallbacksUnmistakably a machine. This is the sound people mean by "robotic".
Concatenative (unit selection)Stitches together small fragments of recorded human speech.Microsoft David / Zira Desktop, Apple "Compact" voicesRecognisably human phonemes with audible joins, and flat, samey sentence melody.
NeuralA model generates the waveform, so rhythm and intonation are learned rather than rule-based.Microsoft "Online (Natural)", Google network voices, Apple Siri voices, and commercial services like the ElevenLabs clips aboveThe jump that stops it sounding robotic. Sentence melody, breath and emphasis land in roughly the right places.

Nothing you do with rate and pitch will turn a formant voice into a neural one. Voice selection is the whole ballgame.

2. The fix that costs nothing: choose the voice, do not accept the default

speechSynthesis.getVoices() hands back every voice on the device, and the one flagged default is whatever the operating system set years ago, which is very often a legacy voice. A plugin can rank that list and pick the best one present. The rules that do almost all the work:

The players on this page now do exactly that. The four voices in the hamburger menu are the top four on your device by that ranking, tagged Natural or Online where they qualify. Compare the top entry with the bottom of the full list and the difference is usually larger than anything else discussed on this issue.

3. What "good" actually looks like per platform

Worth knowing before anyone judges quality from a single machine, because the same plugin genuinely sounds different depending on where it runs:

Platform and browserBest free voices usually presentWhat it falls back to
Windows 11 + EdgeMicrosoft Ava, Andrew, Emma, Brian Online (Natural). Neural, free, no key.Microsoft David / Zira Desktop
Windows + Chrome"Google US English" and siblings, network voicesMicrosoft Desktop voices
macOS / iOS SafariSiri voices, and Samantha (Enhanced)Compact and novelty voices
Android ChromeGoogle Speech Services voicesDevice manufacturer engine, sometimes eSpeak
Linux FirefoxWhatever speech-dispatcher exposes, often nothing goodeSpeak, the worst case

The uncomfortable one: Edge exposes far better voices than Chrome does on the identical Windows machine. So "how robotic is it" partly depends on a browser choice the Learn team does not control. That is a real argument for not relying on the browser route alone for the lessons that matter most.

4. Prosody: what is actually controllable, and what is not

5. The short version

Most of the fix is free, and it is one function. Rank the device's voices and pick the best one instead of taking the browser default, speak sentence by sentence, leave pitch alone, and write lesson text that is speakable. On a current Windows or Apple device that alone moves it from "robotic" to "fine", at no cost and with nothing to install on WordPress.org.

The ceiling is still real. No browser voice will match a produced narrator, and an older Android or a Linux box will sound poor whatever we pick. So for a small set of flagship lessons, and for non-English, the pre-generated file route is what gets you "not robotic at all", and it still needs no live API and no billing. That is the same split recommended below.

What this suggests

  • Default to the free browser-native route. No cost, no third-party dependency, no data leaves the device, no privacy review. It ships as a small plugin and is the only route that fits WordPress.org infrastructure without a procurement decision.
  • For high-value lessons and for non-English, allow pre-generated static audio files. Made once, committed like any other asset. Commercial quality, still no live paid API and no billing on wp.org.
  • Text-only lessons are where this helps most, because there is no narration at all today, and text can be corrected in minutes where a video cannot.
  • Keep the live commercial API as a comparison only, not as the thing we install.
  • Urdu answers the English-only versus all-locales question with evidence: the browser route cannot speak it on a typical device, so translated content realistically needs the pre-generated-file route.

What this demo is and is not

This is a demonstration to make the discussion concrete. It is not a proposal to install any service on WordPress.org, which would be a Meta team decision. The commercial clips were produced on a personal free-tier allowance purely to illustrate quality, with the exact service, voices and settings stated above so the comparison is reproducible.