Agent Skillssanjay3290/ai-skills › elevenlabs

elevenlabs

GitHub

利用 ElevenLabs API 将文本或文档转换为高质量音频,支持单语音朗读和双主播播客生成模式。

skills/elevenlabs/SKILL.md sanjay3290/ai-skills

Trigger Scenarios

创建播客 文档旁白 文本转语音

Install

npx skills add sanjay3290/ai-skills --skill elevenlabs -g -y
More Options

Use without installing

npx skills use sanjay3290/ai-skills@elevenlabs

指定 Agent (Claude Code)

npx skills add sanjay3290/ai-skills --skill elevenlabs -a claude-code -g -y

安装 repo 全部 skill

npx skills add sanjay3290/ai-skills --all -g -y

预览 repo 内 skill

npx skills add sanjay3290/ai-skills --list

SKILL.md

Frontmatter
{
    "name": "elevenlabs",
    "license": "Apache-2.0",
    "metadata": {
        "author": "sanjay3290",
        "version": "1.0"
    },
    "description": "Convert documents and text to audio using ElevenLabs text-to-speech.\nUse this skill when the user wants to create a podcast, narrate a document,\nread aloud text, generate audio from a file, or convert text to speech.\n"
}

ElevenLabs - Text-to-Speech & Podcast Skill

Overview

This skill converts text and documents into high-quality audio using ElevenLabs TTS API. It supports two modes: single-voice narration and two-host conversational podcast generation.

When to Use This Skill

Activate when the user mentions:

  • "create podcast", "generate podcast", "podcast from document"
  • "narrate document", "narrate this file", "read aloud"
  • "text to speech", "TTS", "convert to audio"
  • "audio from document", "audio version of"

Setup

Config at skills/elevenlabs/config.json:

{
  "api_key": "your-elevenlabs-api-key",
  "default_voice": "JBFqnCBsd6RMkjVDRZzb",
  "default_model": "eleven_multilingual_v2",
  "podcast_voice1": "JBFqnCBsd6RMkjVDRZzb",
  "podcast_voice2": "EXAVITQu4vr4xnSDxMaL"
}

Only api_key is required. Or set ELEVENLABS_API_KEY env var.

Dependencies: pip install PyPDF2 python-docx (only needed for PDF/DOCX files).

Requires ffmpeg for multi-chunk narration and podcasts.

Commands

List Voices

python skills/elevenlabs/scripts/elevenlabs.py voices
python skills/elevenlabs/scripts/elevenlabs.py voices --json

Use this to find voice IDs for the user.

Single-Voice TTS

# From text
python skills/elevenlabs/scripts/elevenlabs.py tts --text "Hello world" --output ~/Downloads/hello.mp3

# From document
python skills/elevenlabs/scripts/elevenlabs.py tts --file /path/to/doc.pdf --output ~/Downloads/narration.mp3

# With specific voice
python skills/elevenlabs/scripts/elevenlabs.py tts --file doc.md --voice VOICE_ID --output out.mp3

The script handles text extraction, chunking at sentence boundaries (~4000 chars), TTS per chunk with voice continuity, and ffmpeg concatenation automatically.

Podcast Generation

Podcast mode requires a JSON script file with conversation segments:

[
  {"speaker": "host1", "text": "Welcome to our podcast! Today we're diving into..."},
  {"speaker": "host2", "text": "That's right! I found the section on..."},
  {"speaker": "host1", "text": "Let's break that down..."}
]
python skills/elevenlabs/scripts/elevenlabs.py podcast --script /tmp/script.json --voice1 ID1 --voice2 ID2 --output ~/Downloads/podcast.mp3

Podcast Workflow (for Claude)

When the user asks to create a podcast from a document:

  1. Extract the document text:

    python skills/elevenlabs/scripts/extract.py /path/to/document.pdf
    
  2. Generate a two-host conversation script from the extracted text. Follow these guidelines:

    • Write as a natural, engaging discussion between two hosts
    • Host 1 typically leads/introduces topics, Host 2 adds analysis and reactions
    • Start with a brief intro welcoming listeners and stating the topic
    • End with a summary/outro
    • Keep each turn under 3000 characters
    • Vary turn lengths - mix short reactions with longer explanations
    • Use conversational language: "That's a great point", "What I found interesting was..."
    • Reference specific details from the source document
    • Avoid reading the document verbatim - discuss and interpret it
  3. Write the script as a JSON array to a temp file:

    # Write to /tmp/podcast_script.json
    [
      {"speaker": "host1", "text": "Welcome to today's episode..."},
      {"speaker": "host2", "text": "Thanks for having me..."},
      ...
    ]
    
  4. Generate the podcast:

    python skills/elevenlabs/scripts/elevenlabs.py podcast --script /tmp/podcast_script.json --output ~/Downloads/podcast.mp3
    
  5. Clean up the temp script file.

Tips

  • Run voices first to let the user pick voices they like
  • For podcasts, suggest voice pairs with contrasting qualities (e.g., one deep, one bright)
  • Default output to ~/Downloads/ unless the user specifies otherwise
  • For large documents, warn the user about character usage on their ElevenLabs plan

Version History

  • 3619692 Current 2026-07-24 17:33

Same Skill Collection

skills/atlassian/SKILL.md
skills/azure-devops/SKILL.md
skills/gmail/SKILL.md
skills/google-calendar/SKILL.md
skills/google-chat/SKILL.md
skills/google-docs/SKILL.md
skills/google-drive/SKILL.md
skills/google-sheets/SKILL.md
skills/google-slides/SKILL.md
skills/grok-build/SKILL.md
skills/imagen/SKILL.md
skills/jules/SKILL.md
skills/manus/SKILL.md
skills/mssql/SKILL.md
skills/mysql/SKILL.md
skills/notebooklm/SKILL.md
skills/postgres/SKILL.md
skills/apple-container/SKILL.md
skills/google-tts/SKILL.md
skills/telegram/SKILL.md
skills/whatsapp/SKILL.md

Metadata

Files
0
Version
3619692
Hash
791a24f6
Indexed
2026-07-24 17:33

inicio - Wiki
Copyright © 2011-2026 iteam. Current version is 2.155.2. UTC+08:00, 2026-08-22 06:07
浙ICP备14020137号-1 $mapa de visitantes$