Dev.to WebDev 🛠 Dev 👁 0 📖 5 min read

Create an AI Voice Agent for Customer Support

Why a Voice Agent Makes Sense for Support Customer support is shifting from static FAQs to real‑time, conversational experiences. Adding a voice channel lets you: Reduce friction for users who prefer speaking over ty

Why a Voice Agent Makes Sense for Support

Customer support is shifting from static FAQs to real‑time, conversational experiences. Adding a voice channel lets you:

  • Reduce friction for users who prefer speaking over typing.
  • Provide a more personal touch that can boost satisfaction scores.
  • Scale human agents by handling routine queries automatically.

In this guide we’ll build a simple AI voice agent that can answer common support questions, synthesize responses with high‑quality text‑to‑speech (TTS), and even clone a brand‑specific voice. All of this can be done with a few lines of Python and the powerful API from ElevenLabs.

What You’ll Need

Item Reason
Python 3.9+ Core language for the demo.
requests library To call the ElevenLabs REST endpoints.
Flask (or FastAPI) Lightweight web framework for the HTTP endpoint.
OpenAI API key (or any LLM you like) Generates the textual answer to the user's query.
ElevenLabs API key Powers the TTS and voice‑cloning features.

Tip: If you don’t have an ElevenLabs account yet, sign up through the affiliate link above – you’ll get a free quota to experiment.

Setting Up ElevenLabs

  1. Create an account at the link above.
  2. Navigate to the API dashboard and copy your API key.
  3. (Optional) Upload a short voice sample (30 s–2 min) of a brand spokesperson to create a custom voice model. The dashboard will give you a voice_id you can reuse in the code.

The Architecture in a Nutshell

User (phone/VOIP) ──► Speech‑to‑Text (STT) ──► LLM (generate answer) ──► ElevenLabs TTS ──► Audio response

In this article we focus on the LLM → TTS leg. You can plug any STT service (Google Speech, Whisper, etc.) in front of it.

Generating Text with an LLM (Python)

import os
import requests

OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")

def get_answer(question: str) -> str:
    """Call OpenAI's chat completion endpoint to generate a support reply."""
    headers = {
        "Authorization": f"Bearer {OPENAI_API_KEY}",
        "Content-Type": "application/json"
    }
    data = {
        "model": "gpt-4o-mini",
        "messages": [
            {"role": "system", "content": "You are a friendly customer support agent for Acme Corp."},
            {"role": "user", "content": question}
        ],
        "temperature": 0.6
    }
    resp = requests.post("https://api.openai.com/v1/chat/completions", json=data, headers=headers)
    resp.raise_for_status()
    return resp.json()["choices"][0]["message"]["content"]

The function returns a concise, polite answer that we’ll later speak aloud.

Converting Text to Speech with ElevenLabs

ElevenLabs offers a high‑fidelity TTS endpoint that can be called with a simple POST. Below is a wrapper that sends the generated answer and streams back an MP3 file.

def synthesize_speech(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> bytes:
    """
    Convert `text` to audio using ElevenLabs.
    - `voice_id` can be a default voice or a custom cloned voice.
    """
    url = f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}"
    headers = {
        "xi-api-key": ELEVEN_API_KEY,
        "Content-Type": "application/json"
    }
    payload = {
        "text": text,
        "model_id": "eleven_monolingual_v1",
        "voice_settings": {
            "stability": 0.75,
            "similarity_boost": 0.85
        }
    }
    response = requests.post(url, json=payload, headers=headers)
    response.raise_for_status()
    return response.content  # MP3 bytes

Why ElevenLabs? Their neural models sound natural even for long paragraphs, and the cloning workflow lets you keep a consistent brand voice across all channels.

Putting It All Together with Flask

from flask import Flask, request, send_file, jsonify
from io import BytesIO

app = Flask(__name__)

# Replace with your custom voice ID if you created one
CUSTOM_VOICE_ID = "YOUR_CUSTOM_VOICE_ID"

@app.route("/support", methods=["POST"])
def support():
    data = request.get_json()
    question = data.get("question")
    if not question:
        return jsonify({"error": "Missing 'question' field"}), 400

    # 1️⃣ Generate answer
    answer = get_answer(question)

    # 2️⃣ Convert answer to speech
    audio_bytes = synthesize_speech(answer, voice_id=CUSTOM_VOICE_ID)

    # 3️⃣ Return audio stream
    return send_file(
        BytesIO(audio_bytes),
        mimetype="audio/mpeg",
        as_attachment=False,
        download_name="response.mp3"
    )

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=5000)

How it works

  1. The client (your phone app, a web widget, or a Twilio webhook) sends a JSON payload: {"question": "How do I reset my password?"}.
  2. The endpoint calls get_answer → synthesize_speech.
  3. The MP3 stream is sent back, ready for playback.

You can easily adapt this to WebSockets for real‑time streaming or integrate it with a telephony provider like Twilio or Vonage.

Voice Cloning in Action

If you uploaded a brand voice sample earlier, you’ll have a voice_id that looks something like e8b3a6c9-.... Use it in the synthesize_speech call:

CUSTOM_VOICE_ID = "e8b3a6c9-1234-5678-90ab-cdef12345678"
audio = synthesize_speech("Welcome back! How can I help you today?", voice_id=CUSTOM_VOICE_ID)

The result will sound exactly like the person in the sample, preserving tone, pacing, and pronunciation quirks. This is perfect for keeping the voice of a beloved support mascot alive across all channels.

Testing the Agent Locally

# Start the Flask server
python app.py

# In another terminal, send a test request
curl -X POST http://localhost:5000/support \
     -H "Content-Type: application/json" \
     -d '{"question":"What is your refund policy?"}' \
     --output response.mp3

Open response.mp3 with any media player – you should hear a natural‑sounding answer spoken in your chosen voice.

Next Steps

  • Add STT: Plug Whisper or Google Speech‑to‑Text so users can speak their question.
  • Cache audio: Frequently asked questions can be pre���generated and stored to reduce latency.
  • Add fallback: If the LLM confidence is low, route the call to a human agent.
  • Analytics: Log question‑answer pairs to continuously improve the knowledge base.

Wrap‑Up

Building a voice‑first support agent doesn’t require a PhD in signal processing. With a few API calls to a state‑of‑the‑art TTS service like ElevenLabs, you can deliver a polished, brand‑consistent voice experience that scales alongside your product.

Ready to give your customers a voice they’ll love? Grab your free ElevenLabs credits via the link above, spin up the Flask example, and start experimenting with voice cloning today! 🚀

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.