Create an AI Voice Agent for Customer Support
Why a Voice Agent Makes Sense for Support Customer support is shifting from static FAQs to real‑time, conversational experiences. Adding a voice channel lets you: Reduce friction for users who prefer speaking over ty
Why a Voice Agent Makes Sense for Support
Customer support is shifting from static FAQs to real‑time, conversational experiences. Adding a voice channel lets you:
- Reduce friction for users who prefer speaking over typing.
- Provide a more personal touch that can boost satisfaction scores.
- Scale human agents by handling routine queries automatically.
In this guide we’ll build a simple AI voice agent that can answer common support questions, synthesize responses with high‑quality text‑to‑speech (TTS), and even clone a brand‑specific voice. All of this can be done with a few lines of Python and the powerful API from ElevenLabs.
What You’ll Need
| Item | Reason |
|---|---|
| Python 3.9+ | Core language for the demo. |
requests library |
To call the ElevenLabs REST endpoints. |
Flask (or FastAPI) |
Lightweight web framework for the HTTP endpoint. |
| OpenAI API key (or any LLM you like) | Generates the textual answer to the user's query. |
| ElevenLabs API key | Powers the TTS and voice‑cloning features. |
Tip: If you don’t have an ElevenLabs account yet, sign up through the affiliate link above – you’ll get a free quota to experiment.
Setting Up ElevenLabs
- Create an account at the link above.
- Navigate to the API dashboard and copy your API key.
- (Optional) Upload a short voice sample (30 s–2 min) of a brand spokesperson to create a custom voice model. The dashboard will give you a
voice_idyou can reuse in the code.
The Architecture in a Nutshell
User (phone/VOIP) ──► Speech‑to‑Text (STT) ──► LLM (generate answer) ──► ElevenLabs TTS ──► Audio response
In this article we focus on the LLM → TTS leg. You can plug any STT service (Google Speech, Whisper, etc.) in front of it.
Generating Text with an LLM (Python)
import os
import requests
OPENAI_API_KEY = os.getenv("OPENAI_API_KEY")
ELEVEN_API_KEY = os.getenv("ELEVEN_API_KEY")
def get_answer(question: str) -> str:
"""Call OpenAI's chat completion endpoint to generate a support reply."""
headers = {
"Authorization": f"Bearer {OPENAI_API_KEY}",
"Content-Type": "application/json"
}
data = {
"model": "gpt-4o-mini",
"messages": [
{"role": "system", "content": "You are a friendly customer support agent for Acme Corp."},
{"role": "user", "content": question}
],
"temperature": 0.6
}
resp = requests.post("https://api.openai.com/v1/chat/completions", json=data, headers=headers)
resp.raise_for_status()
return resp.json()["choices"][0]["message"]["content"]
The function returns a concise, polite answer that we’ll later speak aloud.
Converting Text to Speech with ElevenLabs
ElevenLabs offers a high‑fidelity TTS endpoint that can be called with a simple POST. Below is a wrapper that sends the generated answer and streams back an MP3 file.
def synthesize_speech(text: str, voice_id: str = "EXAVITQu4vr4xnSDxMaL") -> bytes:
"""
Convert `text` to audio using ElevenLabs.
- `voice_id` can be a default voice or a custom cloned voice.
"""
url = f"https://api.elevenlabs.io/v1/text-to-speech/{voice_id}"
headers = {
"xi-api-key": ELEVEN_API_KEY,
"Content-Type": "application/json"
}
payload = {
"text": text,
"model_id": "eleven_monolingual_v1",
"voice_settings": {
"stability": 0.75,
"similarity_boost": 0.85
}
}
response = requests.post(url, json=payload, headers=headers)
response.raise_for_status()
return response.content # MP3 bytes
Why ElevenLabs? Their neural models sound natural even for long paragraphs, and the cloning workflow lets you keep a consistent brand voice across all channels.
Putting It All Together with Flask
from flask import Flask, request, send_file, jsonify
from io import BytesIO
app = Flask(__name__)
# Replace with your custom voice ID if you created one
CUSTOM_VOICE_ID = "YOUR_CUSTOM_VOICE_ID"
@app.route("/support", methods=["POST"])
def support():
data = request.get_json()
question = data.get("question")
if not question:
return jsonify({"error": "Missing 'question' field"}), 400
# 1️⃣ Generate answer
answer = get_answer(question)
# 2️⃣ Convert answer to speech
audio_bytes = synthesize_speech(answer, voice_id=CUSTOM_VOICE_ID)
# 3️⃣ Return audio stream
return send_file(
BytesIO(audio_bytes),
mimetype="audio/mpeg",
as_attachment=False,
download_name="response.mp3"
)
if __name__ == "__main__":
app.run(host="0.0.0.0", port=5000)
How it works
- The client (your phone app, a web widget, or a Twilio webhook) sends a JSON payload:
{"question": "How do I reset my password?"}. - The endpoint calls
get_answer→synthesize_speech. - The MP3 stream is sent back, ready for playback.
You can easily adapt this to WebSockets for real‑time streaming or integrate it with a telephony provider like Twilio or Vonage.
Voice Cloning in Action
If you uploaded a brand voice sample earlier, you’ll have a voice_id that looks something like e8b3a6c9-.... Use it in the synthesize_speech call:
CUSTOM_VOICE_ID = "e8b3a6c9-1234-5678-90ab-cdef12345678"
audio = synthesize_speech("Welcome back! How can I help you today?", voice_id=CUSTOM_VOICE_ID)
The result will sound exactly like the person in the sample, preserving tone, pacing, and pronunciation quirks. This is perfect for keeping the voice of a beloved support mascot alive across all channels.
Testing the Agent Locally
# Start the Flask server
python app.py
# In another terminal, send a test request
curl -X POST http://localhost:5000/support \
-H "Content-Type: application/json" \
-d '{"question":"What is your refund policy?"}' \
--output response.mp3
Open response.mp3 with any media player – you should hear a natural‑sounding answer spoken in your chosen voice.
Next Steps
- Add STT: Plug Whisper or Google Speech‑to‑Text so users can speak their question.
- Cache audio: Frequently asked questions can be pre���generated and stored to reduce latency.
- Add fallback: If the LLM confidence is low, route the call to a human agent.
- Analytics: Log question‑answer pairs to continuously improve the knowledge base.
Wrap‑Up
Building a voice‑first support agent doesn’t require a PhD in signal processing. With a few API calls to a state‑of‑the‑art TTS service like ElevenLabs, you can deliver a polished, brand‑consistent voice experience that scales alongside your product.
Ready to give your customers a voice they’ll love? Grab your free ElevenLabs credits via the link above, spin up the Flask example, and start experimenting with voice cloning today! 🚀
Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.