Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 3 min read

My Voice Agent Replied in 8 Seconds. Two Settings, Not the Model.

I run OpenClaw as my main agent harness, with the voice layer on Discord. When I first turned voice on, I said hello and waited about eight seconds for a reply. Eight seconds is long enough to assume it doesn't work. It

I run OpenClaw as my main agent harness, with the voice layer on Discord. When I first turned voice on, I said hello and waited about eight seconds for a reply. Eight seconds is long enough to assume it doesn't work.

It works. The delay was two configuration defaults. Neither was the model.

OpenClaw is the open-source agent harness I run my workflow on; the voice layer sits on Discord. The config keys are OpenClaw's, but the cause β€” turn-taking defaults β€” is generic.

Two defaults caused most of the delay

A fixed silence wait. captureSilenceGraceMs defaults to 2000 ms. It waits two full seconds after you stop talking before it decides your turn is over. Normal conversation leaves a few hundred milliseconds. Two seconds reads as a pause.

A full-agent round-trip on every turn. The realtime voice model consulted the entire agent before it spoke. I'd say "hello"; OpenClaw would run the whole agent β€” tools, memory, the lot β€” and only then would the voice layer talk. That round-trip is the larger half of the delay β€” the same failure I wrote about from the theory side: the pause lives in the turn-taking, not the LLM.

Both are turn-taking settings, not model quality.

The config

Everything sits under channels.discord.voice in ~/.openclaw/openclaw.json. Before β†’ after:

Setting Before (default) After Why
captureSilenceGraceMs 2000 800 Wait a beat, not two seconds
realtime.consultPolicy always auto Voice layer handles the conversational filler; call the agent only when the turn needs it
realtime.bargeIn (on) true Pinned β€” cut it off mid-sentence
realtime.minBargeInAudioEndMs 250 0 Interrupt with no 250ms tail gate
realtime.instructions ends …conversational. Never open with filler. ends …conversational. Lets it use a short backchannel ("one sec") while an agent consult runs
model (agent behind voice) openai/gpt-5.6-terra deepseek/deepseek-v4-flash Cheaper, faster reasoning behind the realtime front end

The one that mattered most was consultPolicy: "auto". The realtime model (gpt-realtime-2.1) is only the voice layer β€” turn-taking, barge-in, playback. The reasoning comes from the agent, which the realtime model calls as a tool. On "always", every turn called the agent before producing audio. On "auto", the voice layer handles the conversational filler itself and calls the agent only when the turn needs it.

The block, after:

"channels": {
  "discord": {
    "voice": {
      "enabled": true,
      "mode": "agent-proxy",
      "captureSilenceGraceMs": 800,
      "realtime": {
        "provider": "openai",
        "model": "gpt-realtime-2.1",
        "speakerVoice": "cedar",
        "providers": {
          "openai": { "apiKey": "***}" }
        },
        "instructions": "Always speak English unless the user is clearly speaking another language to you. Keep spoken replies short, natural, and conversational.",
        "consultPolicy": "auto",
        "bargeIn": true,
        "minBargeInAudioEndMs": 0
      },
      "model": "deepseek/deepseek-v4-flash"
    }
  }
}

You need an API key for the realtime provider. OpenAI here; Grok if you want to price that side.

Result

The realtime model answers first, then reports what it called. The pause is gone and the back-and-forth is responsive.

Cost

One session, over the drive from home to the office: $3.25, mostly input β€” the base listening. Roughly thirty cents a minute on the realtime model while you talk.

Known issues

  • Over Bluetooth, reply audio cuts out in a way music doesn't. The link is less stable than it should be.
  • Long dictation and quick back-and-forth are different modes. This setup is good at the back-and-forth; for long dictation I use a separate dictation tool.

If voice on your agent feels slow, check the silence grace and whether the voice layer consults the full agent every turn before you blame the model. If you've hit a different turn-taking bottleneck, I'd like to know which one.

More on this

Originally published at engineering.kenmazaika.com.

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.