โ† Back to Blog

September 22, 2026 ยท By JayyRedd

Your Chat Just Became a Video Studio โ€” The ElevenLabs MCP Walkthrough

๐ŸŽฌ Watch the full walkthrough: Your Chat Just Became a Video Studio โ€” the install, the four prompts, and the credit maths, on camera.

If you still copy a script out of ChatGPT, paste it into a text-to-speech site, then drag the WAV into CapCut โ€” you're doing 2014 work with 2026 tools.

On September 14th, ElevenLabs turned their MCP connector into a full production studio. One login. Voiceovers, transcripts, dubbing, music, sound effects, images, video. More than fifty models, all inside the chat app you already have open โ€” ChatGPT, Claude, Cursor, Grok, take your pick.

Quick note before we go further, same as I say on camera: the face and voice in the video are my AI avatar. I write every word, the research and the opinions are mine, the delivery is generated. Felt right to say that out loud on a piece about AI production tools.

The thing it actually attacks isn't the idea

Writing the idea was never the hard part.

The hour you lose is the one where you walk a single script through six different apps to get one finished forty-five second piece. Script in one place, voice in another, music in a third, captions in a fourth, timeline in CapCut, dub in another tab, export, miss a sync, start over.

That's eight steps. The new path is three: brief the assistant you already live in, first-cut assets appear, open Studio and fix two things.

The model was never the bottleneck. The handoff was.

MCP, decoded once

MCP stands for Model Context Protocol. Anthropic open-sourced it in November 2024, and the nickname that stuck is "USB-C for AI."

Here's the whole idea: instead of every chat app building a custom integration with every tool on earth, a company publishes one MCP server, and every chat app plugs into the same socket.

Three pieces:

  • Host โ€” the chat app you already use
  • Client โ€” the connector inside that app
  • Server โ€” ElevenLabs' tools at their endpoint

You type in plain English, the assistant calls ElevenLabs, files come back. That's it. I drop the acronym after that.

This is chapter three, and the story matters

April 2025 โ€” local server, developers only. You pasted an API key into Claude Desktop or Cursor. Powerful, fiddly, clearly built for engineers.

August 2026 โ€” ElevenLabs started hosting it themselves. OAuth replaced API keys. It landed in Claude's connector directory. But the point of that release was agent operations: creating voice agents, rewriting prompts, reading call transcripts, estimating cost before a change ships.

September 2026 โ€” same connector, now also a creative factory.

First they let developers talk to the API from chat. Then they let teams run phone agents from chat. Now they let you produce the whole piece from chat.

And here's the part people are missing: if you connected ElevenLabs back in August for the agent stuff, the creative tools showed up on that same install. Nothing new to set up.

Installing it is a minute, not a tutorial series

  • ChatGPT โ€” plugin directory. Search ElevenLabs, connect, sign in.
  • Claude โ€” connectors directory, same flow.
  • Cursor โ€” marketplace.
  • Grok Bot โ€” plugin directory. Codex and OpenClaw are supported too.
  • Claude Code โ€” one command:
claude mcp add --transport http elevenlabs https://api.elevenlabs.io/v1/mcp

Then type /mcp to sign in. Hermes has its own one-liner.

Three things worth knowing while you're in there:

  1. It's OAuth, not an API key. You're never pasting a secret into a client, and you can revoke from either side any time.
  2. You pick data residency at connection time โ€” global, EU, India, or Singapore.
  3. Workspace admins gate the tools. In Claude-class clients you can make individual tools ask for confirmation before they run, and a tool an admin disables cannot be switched back on by a user.

What to actually type

Start with the creative side. Ask for a voiceover in plain language โ€” name the voice, describe the delivery โ€” and the audio generates inside the conversation. No export step, no second tab.

Hand it a recording and ask for a transcript. That's Scribe: speaker labels and timestamps across 99 languages. From that same transcript you can go straight to a script, to captions, or to a dub without leaving the thread.

Ask for the track in another language and dubbing runs through the same connector, built to keep the original speaker's voice and delivery rather than replacing it with a stranger.

Ask for music and you get a track in whatever genre, with or without vocals. Ask for a sound effect and you get variations to pick from. Ask for an image, edit that image, animate it into video, add lipsync โ€” all in sequence, all in the same conversation.

Four prompts you can steal

One โ€” the brand short.

Write a 20-second script for a custom fitting video. Generate the voiceover in a warm male American voice. Make a clean studio image of a tailored navy suit. Animate it with lipsync. Add a low cinematic bed and subtle cloth SFX.

Two โ€” the one people screenshot.

Here's a 4-minute client testimonial from my phone. Transcribe it, clean the ums, write a 30-second highlight, voice it in my brand voice, dub a Spanish version that keeps the same delivery, and make a caption-ready clip.

Three โ€” the one that pays the bills.

Same install, no new setup โ€” spin up a receptionist that answers after hours, answers questions from my site, books the appointment, and texts the confirmation.

Four โ€” a safety habit, not a demo.

Duplicate my production agent first. Change the voice on the copy. Estimate the cost. Do not touch live until I say so.

That last one isn't a flex. Generating content is low-risk. Writing a live production phone agent from a chat window is not, and a confirmation prompt is not a staging environment.

Two product names everybody mixes up

ElevenCreative is the workspace. Everything you generate lands there.

Studio is the timeline editor inside that workspace โ€” tracks for voiceover, music, SFX, video, captions. You can regenerate one sentence, or one word, and lock a section once it's good.

ElevenLabs' own framing is that the conversation gets you to a strong first version fast, and Studio is where you take precise control of the final cut. I'd put it harder than that:

Chat is your first draft factory. Studio is where taste lives. Don't let anybody sell you "one prompt, finished ad, posted."

Money, because the comments are undefeated

There's no separate fee for the connector. It spends your normal ElevenLabs credits.

Plan Price Credits Note
Free $0 10k Non-commercial
Starter $6 30k First commercial licence
Creator $22 ~121k Professional voice cloning
Pro $99 600k Higher-quality API audio

Confirm these live on the pricing page before you quote them anywhere โ€” plans move.

The thing nobody tells you is that credit burn is wildly unequal across media. Published ballparks:

  • Transcription โ€” ~330 credits/minute
  • Sound effect โ€” ~200 credits/generation
  • Music โ€” ~900 credits/minute
  • Automatic dubbing โ€” ~2,000 credits/minute

So one lazy "make me the whole video" prompt can fire several expensive operations at once. Measure cost per finished asset, not per prompt โ€” every retake bills again.

I'll put a real number on it. The metaphor b-roll in the video โ€” eight silent clips โ€” was generated through this same connector on Veo 3.1 Lite at 2,424 credits each. Two came back with garbled text baked onto a key and a cable connector, so they got reshot. Ten submissions, eight usable, 24,240 credits, about $5.33. That's the honest shape of it: you budget for rejects.

And the trap: the free plan is not commercial. If you're publishing or selling the work, you want Starter or above.

What it does not replace

  • Taste, brand voice, and a real edit
  • A guarantee the routed image/video model is the one you'd have picked
  • CapCut or Premiere for a tight, punchy final cut
  • Free unlimited generation โ€” credits are the meter
  • A safe yes on live agent edits

The honest frame

It isn't that AI is going to make your content. It's that AI is going to stop making you the production assistant on your own content.

So here's the homework, and it isn't a demo: go make the video you already owed somebody this week. Make it the old way, make it this way, and time both.

That's the only review that matters.


Links worth having open