Skip to content
Skillv1.0.0

elevenlabs-stt

ElevenLabs speech-to-text with Scribe models and forced alignment via inference.sh CLI. Models: Scribe v1/v2 (98%+ accuracy, 90+ languages). Capabilities: transcription, speaker diarization, audio eve

by skills-101(0) 0 installs
Free
Sign in to install

Free account. Installing gives you the manifest plus copy-paste snippets.

See reviews

About

Imported from skills-101/superpowers (tools/audio/elevenlabs-stt/SKILL.md). Install upstream with npx skills add skills-101/superpowers --skill elevenlabs-stt. Copyright stays with the author.

Install the belt CLI skill: npx skills add belt-sh/cli

ElevenLabs Speech-to-Text

High-accuracy transcription with Scribe models via inference.sh CLI.

ElevenLabs STT

Quick Start

Requires inference.sh CLI (belt). Install instructions

belt login

# Transcribe audio
belt app run elevenlabs/stt --input '{"audio": "https://audio.mp3"}'

Available Models

Model ID Best For
Scribe v2 scribe_v2 Latest, highest accuracy (default)
Scribe v1 scribe_v1 Stable, proven
  • 98%+ transcription accuracy
  • 90+ languages with auto-detection

Examples

Basic Transcription

belt app run elevenlabs/stt --input '{"audio": "https://meeting-recording.mp3"}'

With Speaker Identification

belt app run elevenlabs/stt --input '{
  "audio": "https://meeting.mp3",
  "diarize": true
}'

Audio Event Tagging

Detect laughter, applause, music, and other non-speech events:

belt app run elevenlabs/stt --input '{
  "audio": "https://podcast.mp3",
  "tag_audio_events": true
}'

Specify Language

belt app run elevenlabs/stt --input '{
  "audio": "https://spanish-audio.mp3",
  "language_code": "spa"
}'

Full Options

belt app run elevenlabs/stt --input '{
  "audio": "https://conference.mp3",
  "model": "scribe_v2",
  "diarize": true,
  "tag_audio_events": true,
  "language_code": "eng"
}'

Forced Alignment

Get precise word-level and character-level timestamps by aligning known text to audio. Useful for subtitles, lip-sync, and karaoke.

belt app run elevenlabs/forced-alignment --input '{
  "audio": "https://narration.mp3",
  "text": "This is the exact text spoken in the audio file."
}'

Output Format

{
  "words": [
    {"text": "This", "start": 0.0, "end": 0.3},
    {"text": "is", "start": 0.35, "end": 0.5},
    {"text": "the", "start": 0.55, "end": 0.65}
  ],
  "text": "This is the exact text spoken in the audio file."
}

Forced Alignment Use Cases

  • Subtitles: Precise timing for video captions
  • Lip-sync: Align audio to animated characters
  • Karaoke: Word-by-word timing for lyrics
  • Accessibility: Synchronized transcripts

Workflow: Video Subtitles

# 1. Transcribe video audio
belt app run elevenlabs/stt --input '{
  "audio": "https://video.mp4",
  "diarize": true
}' > transcript.json

# 2. Use transcript for captions
belt app run infsh/caption-videos --input '{
  "video_url": "https://video.mp4",
  "captions": "<transcript-from-step-1>"
}'

Supported Languages

90+ languages including: English, Spanish, French, German, Italian, Portuguese, Chinese, Japanese, Korean, Arabic, Hindi, Russian, Turkish, Dutch, Swedish, and many more. Leave language_code empty for automatic detection.

Use Cases

  • Meetings: Transcribe recordings with speaker identification
  • Podcasts: Generate transcripts with audio event tags
  • Subtitles: Create timed captions for videos
  • Research: Interview transcription with diarization
  • Accessibility: Make audio content searchable and accessible
  • Lip-sync: Forced alignment for animation timing

Related Skills

# ElevenLabs TTS (reverse direction)
npx skills add inference-sh/skills@elevenlabs-tts

# ElevenLabs dubbing (translate audio)
npx skills add inference-sh/skills@elevenlabs-dubbing

# Other STT models (Whisper)
npx skills add inference-sh/skills@speech-to-text

# Full platform skill (all apps)
npx skills add inference-sh/skills@infsh-cli

Browse all audio apps: belt app list --category audio

Use it

Copy one of these into your project. Installing also returns the manifest and these snippets.

yaml
targets:
  - https://api.opensmartroute.ai/api/v1/registry/skills-101-superpowers-elevenlabs-stt/manifest   # or paste the manifest below

Manifest

An Open Capability Manifest: the router reads it to know what this does, what it costs and when to pick it.

skills-101-superpowers-elevenlabs-stt.ocm.jsonjson
{
  "ocm": "1",
  "id": "skills-101-superpowers-elevenlabs-stt",
  "kind": "skill",
  "name": "elevenlabs-stt",
  "description": "ElevenLabs speech-to-text with Scribe models and forced alignment via inference.sh CLI. Models: Scribe v1/v2 (98%+ accuracy, 90+ languages). Capabilities: transcription, speaker diarization, audio event tagging, word-level timestamps, forced alignment, subtitle generation. Use for: meeting transcription, subtitles, podcast transcripts, lip-sync timing, karaoke. Triggers: elevenlabs stt, elevenlabs transcription, scribe, elevenlabs speech to text, forced alignment, word alignment, subtitle timing, diarization, speaker identification, audio event detection, eleven labs transcribe",
  "publisher": "skills-101",
  "version": "1.0.0",
  "capabilities": {
    "domains": [
      "math"
    ],
    "tags": [
      "skill-md",
      "skills-sh"
    ],
    "languages": [
      "en"
    ]
  },
  "quality_prior": 0.6,
  "examples": [
    "ElevenLabs speech-to-text with Scribe models and forced alignment via inference.sh CLI. Models: Scribe v1/v2 (98%+ accuracy, 90+ languages). Capabilities: transcription, speaker diarization, audio event tagging, word-level timestamps, forced alignment, subtitle generation. Use for: meeting transcription, subtitles, podcast transcripts, lip-sync timing, karaoke. Triggers: elevenlabs stt, elevenlabs transcription, scribe, elevenlabs speech to text, forced alignment, word alignment, subtitle timing, diarization, speaker identification, audio event detection, eleven labs transcribe"
  ],
  "primary": false,
  "metadata": {
    "source": {
      "provider": "skills.sh",
      "repository": "https://github.com/skills-101/superpowers",
      "path": "tools/audio/elevenlabs-stt/SKILL.md",
      "ref": "HEAD",
      "url": "https://github.com/skills-101/superpowers/blob/HEAD/tools/audio/elevenlabs-stt/SKILL.md",
      "key": "skills-101/superpowers/tools/audio/elevenlabs-stt/SKILL.md"
    },
    "allowed_tools": [
      "Bash(belt",
      "*)"
    ]
  },
  "instructions": "> **Install the belt CLI skill:** `npx skills add belt-sh/cli`\n\n# ElevenLabs Speech-to-Text\n\nHigh-accuracy transcription with Scribe models via [inference.sh](https://inference.sh) CLI.\n\n![ElevenLabs STT](https://cloud.inference.sh/u/4mg21r6ta37mpaz6ktzwtt8krr/01jz025e88nkvw55at1rqtj5t8.png)\n\n## Quick Start\n\n> Requires inference.sh CLI (`belt`). [Install instructions](https://raw.githubusercontent.com/inference-sh/skills/refs/heads/main/cli-install.md)\n\n```bash\nbelt login\n\n# Transcribe audio\nbelt app run elevenlabs/stt --input '{\"audio\": \"https://audio.mp3\"}'\n```\n\n\n## Available Models\n\n| Model",
  "cost": {
    "context_tokens": 958
  }
}

Fetch it by URL: GET /api/v1/registry/skills-101-superpowers-elevenlabs-stt/manifest?version=1.0.0

Reviews

Star ratings from people who tried it. One review per account; edit yours any time.

No reviews yet. Install it, try it, and be the first to rate it.