A good chunk of your audience will never read what you write.

Not because it's bad. Because they're driving. Or at the gym. Or walking the dog. Or they're the kind of person who'd happily listen to you for twenty minutes but hasn't read past the second paragraph of an email since 2019.

For years the answer to that was "start a podcast." Which meant a microphone, a quiet room, an editing workflow, a hosting account, and an extra half day a week you didn't have. So most people didn't, and that whole slice of the audience just never heard from them.

That changed a while ago, but last week it got good enough that I'd stop thinking of it as a gimmick. ElevenLabs shipped its v4 voice models, in a standard version and a faster Turbo version, with support for more than ninety languages and something they call inline performance direction. Inworld released a realtime voice model the same week that gives you control over tone, pacing and even laughs.

Inline direction is the piece that matters. It means you can mark up a script the way a director marks up a page. Slow down here. Lighter here. Let this land. The flat, slightly too perfect read that made AI audio sound like an airport announcement is now something you can steer.

So this week we're building an audio edition. Your written content, turned into something people can listen to, in your voice, without you recording a thing.

What you're actually building

Let me be specific so nobody builds the wrong thing.

You're not building a podcast in the "two people chatting for an hour" sense. You're building a spoken version of content you already write, published on a schedule, in a feed people can subscribe to.

If you write a weekly newsletter, it's the audio version of that newsletter. If you write blog posts, it's the audio version of your best ones. If you write long LinkedIn posts, it's a short weekly audio roundup of them.

The bar to clear isn't "sounds like a professional podcast." It's "sounds like you reading your own stuff on a good day." That bar is very clearable now.

Step one: rewrite for the ear

This is the step everyone skips and it's the reason most AI audio sounds bad. Not the voice. The script.

Writing for the eye and writing for the ear are different jobs. Readers can skim, glance back, see a list and take it in at once. Listeners can't. They get each word once, in order, at your speed.

Here's what breaks when you read written content out loud:

Links. "Click here" means nothing in audio. Neither does a URL read aloud.

Lists and tables. Six bullets read in a row turn into mush. A table is unlistenable.

Long sentences with nested clauses. Fine on the page, where your eye can hold the structure. Out loud, people lose the thread by the second comma.

Visual references. "As you can see below." They can't.

Abbreviations and symbols. "$4M" might get read as "four M." "Q4" might come out fine or might not. Percent signs, slashes and ampersands are a coin flip.

So before anything goes to a voice model, it goes through a rewrite. Here's the prompt I use:

❝

Rewrite the piece below as a script to be read aloud by its author. Keep the author's voice, opinions and humor. Keep every substantive point. Then make these changes:

Replace links with a short spoken pointer, like "the link's in the written version."

Turn any list of more than three items into flowing sentences, or say how many items there are before listing them.

Break any sentence over 25 words into two.

Remove references to anything visual.

Spell out numbers, currencies, percentages and abbreviations the way a person would say them out loud.

Add a one sentence spoken intro and a one sentence spoken outro.

Don't add new ideas. Don't make it more formal.

Then read the output yourself, quickly, out loud. Anywhere you trip over a phrase, so will the voice model. Fix those spots by hand. Budget ten or fifteen minutes a piece. Skip it and people bail on your audio halfway through.

Step two: direct the read

Now you've got a script that works for the ear. Before you hand it to the voice, mark it up.

The newer voice models let you add performance notes in the text itself. The exact syntax varies by tool, so check the documentation for the one you're using, but the kinds of direction you can give are similar across them: pace, emphasis, pauses, tone shifts, warmth.

Don't overdo it. Direction on every line sounds like a bad audiobook. Here's where it earns its keep:

The opening. A touch slower and warmer than the rest. People decide whether to keep listening in about ten seconds.

The main point. Wherever the one idea you want people to remember lives, slow down and add a pause after it.

Jokes and asides. Lighter, a little faster. If you write with dry humor, this is where flat AI delivery used to kill it.

Transitions between sections. A short pause. Listeners need a beat to know the topic changed. On the page, a heading does that job. In audio, silence does.

The call to action. Clear, unhurried, said once. Not shouted.

A useful trick is to ask a model to do a first pass of markup for you, then edit it down. Something like: "Add pacing and emphasis notes to this script, using no more than one note per paragraph, focused on the main point, the humor and the transitions." Then delete about half of what it adds. You want a light touch.

Step three: get the voice right

You have two options here, and the right one depends on how you feel about it.

Clone your own voice. Most of the serious voice tools will build a model of your voice from a clean recording. Spend real time on that recording. Quiet room, decent microphone, consistent distance from the mic, talking the way you actually talk, not in your phone voice. The model inherits everything: your pace, your quirks, and any echo in your kitchen. Record more than the minimum the tool asks for. More clean audio almost always means a better clone.

Use a stock voice and say so. Some people would rather not clone themselves, and that's a completely reasonable call. Pick a stock voice that fits your brand and introduce the audio edition honestly: "This is the audio version of this week's issue, narrated by an AI voice so I can get it to you the same day it publishes." Nobody minds. What people mind is being fooled.

If you do clone your voice, I'd still say so. A short line in the intro or the show notes. It's the same rule I'd apply to video avatars. The words are yours, you wrote them and approved them, and the only thing automated is the reading. Saying so costs you nothing. Getting caught not saying so costs you the reason people subscribed.

FROM THE AI NEWSROOM

The AI Workflow Blueprint

The exact systems behind everything in this issue. The audit sheets, the routing logic, the templates and the review cadences, built out step by step so you can copy them straight into your own stack. One time, forty seven dollars.

Get the Blueprint for $47

.....

Step four: build the pipeline

Here's how the whole thing runs without you touching it every week.

The trigger is your piece going live. A newsletter publishing, a blog post going out, whatever your main written channel is.

A Make.com scenario picks it up, pulls the text, and sends it through the rewrite prompt. The script lands in a review spot you'll actually check, a Notion database, a Google Doc, a Slack message. You read it, fix anything that trips you up, and approve it.

On approval, the script goes to your voice tool through its API, the audio comes back, and the file gets uploaded to your podcast host as a new episode with the title, a short description and a link back to the written version.

If you publish on beehiiv, it can host a podcast alongside your newsletter, which saves you running a separate hosting account and keeps everything under one roof. Otherwise any standard podcast host works fine. What matters is that the feed reaches the big podcast apps, because that's where listeners actually live.

Total weekly time once it's built: the five or ten minutes you spend reviewing the script. Everything else runs on its own.

Step five: cut it into pieces

An audio edition gives you something most written creators never have: a library of short, clean audio clips.

Every episode has two or three moments worth pulling out. The main point. A story. A line that made you smile when you wrote it.

Cut those into thirty to sixty second clips, put captions over a simple branded background or a still image, and you've got social content that doesn't look like everyone else's text post. People stop scrolling for audiograms and captioned clips in a way they don't for a quote card, especially on platforms that autoplay.

Schedule them through Buffer across the week following each episode. One piece of writing now becomes a written edition, an audio episode, and three or four clips. You wrote it once.

The languages question

Ninety plus languages gets its own section, because almost nobody's talking about it.

If any meaningful share of your audience or your customers speaks something other than English as a first language, you can now publish an audio edition in their language. Translate the script, have a native speaker skim it if you can, and render it in your cloned voice.

I'd start with one language, not ten. Pick the one where you already have customers or obvious demand. Publish it as a separate feed so people can subscribe in the language they want. And be honest about it: a short note that the translation and narration are AI assisted.

For a one person business or a small team, this used to be completely out of reach. Now it's an afternoon of setup and a review step per episode. If you've been eyeing an international market for years, it's worth pricing it out again.

What to watch out for

A few real cautions before you go build this.

Pronunciation. Your company name, your clients' names, industry terms, anything unusual. Listen to the first full render with a sharp ear. Most voice tools have a pronunciation dictionary or let you spell words phonetically. Fix it once and it stays fixed.

Length. A 2,500 word article becomes roughly fifteen to eighteen minutes of audio. That's fine for an engaged subscriber on a commute. If your pieces run much longer, consider an audio edition that covers the main sections and points to the written version for the rest.

Quality drift. Voice models get updated. When your tool ships a new version, re listen to an episode before assuming it sounds the same. Sometimes it's better. Sometimes your carefully tuned voice comes out subtly different.

Rights. If you're quoting other people's writing at length, that's the same issue in audio as on the page. Keep quotes short.

Don't let it replace the real thing entirely. Once a month or once a quarter, record something yourself. A personal note, a reaction to a big week, a story. Real you, real mic. It reminds people there's an actual person behind the feed, which is kind of the whole point.

Why this matters now

Quick word on why I'd do this now and not next year.

Most of your competitors publish in exactly one format. Text. Their audience is limited to people who like to read, at the moments they have time to read.

An audio edition reaches the same people at different moments. The commute, the workout, the drive between job sites, the hour of yard work on Saturday. Those are hours your written content was never going to get. And attention in someone's ears for fifteen minutes is a very different relationship from a skim of a subject line.

The cost of building this collapsed this year. The quality crossed the line from "obviously robotic" to "genuinely pleasant to listen to" about five minutes ago. Cheap and good doesn't stay a secret. Get in before your competitors figure that out.

FROM THE AI NEWSROOM

The AI Business Accelerator

For operators who would rather build it with someone than build it alone. The full stack, installed alongside you, with the decisions made in the room instead of guessed at. Ninety seven dollars.

Join the Accelerator for $97

.....

This week

Pick your single best piece of written content from the last month. The one that got the most replies or the best feedback.

Run it through the rewrite prompt. Read the script out loud once and fix where you stumble. Add light direction to the opening, the main point and the close.

Render it with a stock voice or your own clone, and listen to the whole thing on a walk.

If you'd happily send it to your list, you've got your first episode. If not, you'll know exactly which step to fix, and it'll almost always be the script, not the voice.

Jordan

The AI Newsroom | Practical AI for people with a business to run.