A ready-made pipeline: an audio episode from text, no microphone. Take an article, a chapter or your notes - AI turns them into a podcast script (a two-host dialogue if you like), you synthesize natural narration, add an intro and music, and publish. The voice is synthetic, so no studio required.
What you end up with
- A finished audio file for the episode (single narrator or a two-"host" dialogue).
- A podcast script built from your text: conversational, with an intro and outro.
- An intro/jingle and a mix (voice sitting above the music).
- Show notes (description + timestamps) for the platforms, plus cover art.
- A repeatable pipeline: the next episode takes an hour.
What you'll need
Tools from the library (by name): ChatGPT/Claude (script, show notes), ElevenLabs (natural narration, including multiple voices) or Kokoro TTS/Chatterbox/Fish Audio (OpenAudio) (open-source, local/free), DaVinci Resolve or OpenCut (mix voice + music, export audio), Descript (optional - edit audio "like text", cut the ums), Nano Banana Pro/Canva (cover art). Accounts/money/hardware:
- Free route: ChatGPT/Claude (free), Kokoro/Chatterbox (local, a GPU helps for local synthesis), OpenCut (in the browser). Any PC will do.
- Fast route: ElevenLabs (~$5-22/mo - the most natural voices and multi-voice dialogue), Descript (free tier, then ~$12-24/mo).
Step by step
- Script from text. Tool: ChatGPT/Claude. What to do: turn your material into a conversational audio script (mono or dialogue). Ready-made prompt:
Turn this text into a 5-8 minute podcast episode script: <paste text>.
Format: a dialogue between two hosts (Host A asks the questions, Host B explains).
Structure: short intro (what the episode is about) -> 3-4 topic blocks, conversational ->
wrap-up and outro with a call to subscribe. Lively speech, no bureaucratese, no filler.
Label the lines "A:" and "B:".
How to tell it worked: it reads like a real conversation, not an article being read aloud; there's an intro and an outro; the lines are labeled. 2. Narration. Tool: ElevenLabs (multi-voice - different voices for A and B) or Kokoro TTS/Fish Audio (free/local). What to do: assign a voice to each host, synthesize line by line, download the audio. Tip: speed 1.0, pauses driven by punctuation. How to tell it worked: the voices are clearly distinct, sound even, no mispronounced words (listen back). 3. Cleanup (optional). Tool: Descript. What to do: if you have live inserts or need trimming - edit the audio "like text", cut pauses and filler words. How to tell it worked: no long pauses or "uhh" on the track, the pace is brisk. 4. Mixing. Tool: DaVinci Resolve or OpenCut. What to do: lay down the narration + intro jingle + quiet background music (-18 to -24 dB under the voice), normalize loudness, export to MP3. How to tell it worked: the voice is clearly louder than the music, no volume jumps between lines, and you get an MP3 out. 5. Show notes + cover art. Tool: ChatGPT/Claude (description/timestamps) + Nano Banana Pro/Canva (cover). Ready-made prompt for show notes:
Write show notes for the episode based on this script: <script>.
Give me: 1) a catchy 3-4 sentence description; 2) a list of timestamps by block;
3) 5 tags/keywords for the platforms.
How to tell it worked: you have a description, timestamps and a cover that's readable as a thumbnail. 6. Publishing. What to do: upload the MP3 to a podcast host (it feeds the RSS to Spotify/Apple/Yandex Music) or post it as audio on Telegram/social media. How to tell it worked: the episode is reachable by link/in the app, with cover art and description in place.
Free route vs fast (paid) route
| Step | Free route | Fast (paid) |
|---|---|---|
| Narration | Kokoro / Chatterbox / Fish Audio (local) | ElevenLabs (~$5-22/mo, multi-voice) |
| Cleanup | by hand in an editor | Descript (edit "like text") |
| Mixing | OpenCut / DaVinci Resolve | DaVinci Resolve (same one, free) |
| Cover art | Canva free | Nano Banana Pro / Canva Pro |
The whole episode really can be built for free (Kokoro + OpenCut). Paying for ElevenLabs buys you voice naturalness and easy two-host dialogue - you can hear the difference, but free TTS already sounds decent.
Common problems and fixes
- The voice sounds robotic. Shorten sentences, add punctuation to drive pauses; in ElevenLabs drop stability to ~40-50%; use a newer TTS model.
- Both hosts sound the same. Assign clearly different voices (timbre/gender/pace) and stick to the roles: A asks short questions, B explains.
- Wrong stress or brand pronunciation. Spell the word the way it sounds (transliterate) or split it with hyphens between syllables.
- Music drowns out the voice. Drop the background to -18 to -24 dB, turn on normalization; the voice always stays front and center.
- It sounds like an article read aloud. In step 1, ask for "a lively conversation, a dialogue, no bureaucratese" and cut the filler; add linking questions between blocks.
How much time and money it takes
- Time: the first episode takes 2-4 hours (while you're learning); after that, working from the template, ~1 hour per episode.
- Money: free (Kokoro/Chatterbox + OpenCut) or ~$5-25/mo for ElevenLabs (+ Descript if you want it).
- Honestly: synthetic voice already sounds good, but it doesn't fully replace a host's live charisma yet; on the other hand, the episode comes together with no studio, no microphone and no editor.