Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 11 min read

Phone calls are the boss fight of living in Japan. ๐Ÿ—ผ

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend I've lived alone in Tokyo for four years now, but when my washing machine started leaking, I stared at the building manager's number for

Phone calls are the boss fight of living in Japan. ๐Ÿ—ผ

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

A practice call in Before I Call: a Japanese question with furigana, a word's meaning on hover, then a plain-English explanation of the question

I've lived alone in Tokyo for four years now, but when my washing machine started leaking, I stared at the building manager's number for twenty minutes before I pressed call.

My friend ran into the same thing when she had to cancel her internet contract by phone. The provider's explanation sounded alien to her. The Japanese businesses use on the phone is not the Japanese you use every day, and on a call there's no face to read and no time to look anything up.

So I built Before I Call, an app where you practice the call with an AI before you make the real one. I made it for myself and for friends in the same situation, and sent it to one of them to see if it actually helps. Their reply is further down.

Here it is in action:

What I Built

You describe the call you need to make, or pick one of the examples (home repairs, the clinic, a missed delivery, the city office, dietary requests, lost property, bills). Then you talk to an AI that plays the person on the other end, like a receptionist or your building manager. It speaks polite Japanese, asks one question at a time and waits for your answer. There's no score.

If you get stuck, you don't have to hang up. Explain question opens a helper next to the call with what the question means and a reply you could give. The AI on the call doesn't see it, so the role-play carries on. Slow replay plays the question again at 0.7x speed, and Repeat question asks the AI to say it again. Kanji have furigana, every line has romaji, and known words show their meaning when you hover or tap.

When you say goodbye, the AI hangs up and you get a call card: the whole conversation with readings and meanings, and a list of useful words from your call. You can download it as a PDF and keep it next to you when you make the real call.

It won't make up dates, prices or availability, and it can't book anything for you. It's only for practice.

Setting up a call. You can type or dictate it in English or Japanese, or start from an example. Enhance prompt tidies up rough notes without adding facts you didn't give it.

The setup form: a washing-machine leak typed in English, with Start from an example, Enhance prompt, Dictate situation and Start voice practice

Asking for help in the middle of the call:

The explanation dialog: the question's meaning, a note on ใฉใ†ใ•ใ‚Œใพใ—ใŸใ‹, and a suggested reply with furigana

The word list on the call card:

The call card's useful words: ็ฎก็†ไผš็คพ, ็”ฐไธญ and ๆด—ๆฟฏๆฉŸ with furigana, romaji and English meanings

What my friend said

I sent the link to my friend on WhatsApp and asked them to let me know if it helps. This is what came back:

WhatsApp chat: I share the Before I Call link with my friend. They reply

I didn't expect the PDF to be the part they liked. I made it as a cheat sheet for the real call, but for someone studying for the JLPT it's also a word list from a conversation they actually had.

Demo

Try it at before-i-call.onrender.com. There's no sign-up. Voice practice needs a microphone. Without one, tap Play guided example to watch a recorded call in Japanese or English.

The PDF call card: the situation, each turn as a speech bubble with furigana, romaji and meaning

Code

Before I Call logo

Before I Call

Rehearse the phone call you've been putting off.

A patient AI voice partner for everyday Japanese and English calls, with furigana on every word,
help that never interrupts the conversation, and a call card to take with you.
Use it in the cloud, or free and offline on your own computer with Gemma 4, Whisper and Kokoro.

Try it live

Gemma 4 open weights ElevenLabs Agents DigitalOcean serverless inference Sentry agent tracing Deployed on Render MIT license Free local mode with Whisper, Gemma and Kokoro 87 tests passing Hacktoberfest 2026

Live app ยท How it works ยท Free local mode ยท Architecture ยท Run it locally


Before I Call home page: practice a phone call in Japanese or English, or play a guided example

The problem

A phone call is the hardest everyday conversation in a second language. There is no face to read and no time to look anything up, and the other person speaks at native speed. People who live in Japan put off calling the building manager, the clinic or the delivery company, not because they can't manage the conversation, but because the call itself feels risky.

Phrasebooks help with the first sentence. They don'tโ€ฆ

MIT licensed and built during the challenge. It has 87 tests (80 Python, 7 Node), and the README explains how to run it locally.

How I Built It

The call runs on ElevenLabs Agents

The live call is an ElevenLabs agent. It handles all the audio: the WebRTC connection, speech recognition, working out when you've finished speaking, interruptions, and the voice. I use one agent for every scenario. When a call starts, the app sends that call's prompt, first line, language and voice as overrides, so the same agent can be a Japanese building manager or an English clinic receptionist.

I slowed the agent's speech to 0.85x for learners and capped calls at two minutes, with a daily limit on the server so the credits can't all go in one day. Hanging up uses the built-in end_call tool. Its description tells the model to use it only after your question is dealt with and you've said you don't need anything else, not when you just say thanks. Explain question opens a second, text-only session with a helper prompt, so the explanation never ends up in the call transcript. Slow replay uses ElevenLabs text to speech (Flash v2.5 at 0.7x).

The agent has had 113 conversations so far. ElevenLabs groups them by topic on its own, and the top three match the examples in the app: booking an appointment (27), the washing machine (19) and a missed delivery (17).

ElevenLabs agent dashboard: 113 conversations, with topics grouped as Appointment Booking and Scheduling (27), Washing Machine Issue Resolution (19) and Missed Delivery Redelivery Request (17)

Why I switched to a custom LLM

When I started, the agent used one of the LLMs hosted by ElevenLabs, Qwen3.5-397B-A17B. Its usage came out of the same ElevenLabs credits I got for Hacktoberfest, and it was going through them fast. ElevenLabs lists it at about $0.016 a minute and $0.003 a message, on top of the voice. A custom LLM costs nothing on the ElevenLabs side:

The ElevenLabs LLM picker: hosted models with their prices, including Qwen3.5-397B-A17B at about $0.0159 per minute and $0.0030 per message, and Custom LLM at $0.0000

I wanted those credits for the voice, so I switched the agent to a custom LLM and pointed it at Gemma 4 31B on DigitalOcean, which bills my DigitalOcean account instead. With the per-token prices my tracing code uses for Gemma ($0.18 per million input tokens and $0.50 per million output tokens), the turn in the trace further down cost about $0.0002. The ElevenLabs credits now only pay for the audio.

At first the custom LLM pointed straight at DigitalOcean. Later I put my own server in between so I could trace and check every reply. This is the agent's LLM setting now:

The agent's Custom LLM setting in ElevenLabs: Chat Completions, server URL https://before-i-call.onrender.com/llm/v1/ and model ID gemma-4-31B-it

On each turn, the agent sends the conversation to my FastAPI server on Render, the server forwards it to Gemma 4 31B on DigitalOcean's serverless inference, and the response streams straight back, so the voice can start as soon as the first words arrive.

Architecture: the browser talks to ElevenLabs over WebRTC; ElevenLabs calls the Gemma proxy on Render, which streams Gemma 4 31B from DigitalOcean and sends traces to Sentry

The proxy on Render

The custom LLM URL points at my own server: a FastAPI app that Render runs as one Docker web service. The same service serves the React app, the API and the proxy. The whole setup is a render.yaml Blueprint in the repo (trimmed here):

services:
  - type: web
    name: before-i-call
    runtime: docker
    plan: starter
    healthCheckPath: /api/health
    disk:
      name: usage-data
      mountPath: /var/data
      sizeGB: 1
    envVars:
      - key: LLM_PROXY_KEY
        sync: false
      - key: GRADIENT_MODEL
        value: gemma-4-31B-it
      - key: MAX_CONCURRENT_CALLS
        value: '2'
      - key: MAX_CALL_SECONDS
        value: '120'
      - key: MAX_CALLS_PER_VISITOR_DAY
        value: '3'

Keys are sync: false, so they're set in the Render dashboard and never committed. The limits that protect my credits are plain environment variables: two calls at once, two minutes each, three calls per visitor a day. The 1 GB disk holds the SQLite files for those limits and the analytics, so they survive redeploys.

The proxy is a single route. It accepts only the agent's key, refuses every model except Gemma 4 31B, and passes DigitalOcean's stream through unchanged while it reads each chunk for tracing and the rule checks (trimmed):

@router.post('/llm/v1/chat/completions')
async def chat_completions(request: Request):
    ...
    if not isinstance(body, dict) or body.get('model') != model:
        # The key is never usable for other, more expensive models.
        raise HTTPException(400, f'Only {model} is available.')
    ...
    async def relay():
        async for chunk in response.aiter_bytes():
            stream.feed(chunk)
            yield chunk

    return StreamingResponse(relay(), media_type='text/event-stream',
                             headers={'Cache-Control': 'no-cache', 'X-Accel-Buffering': 'no'})

An always-on web service suits this. Each reply is a streamed response that stays open while Gemma talks. The server reuses one HTTPS client for DigitalOcean, so a turn doesn't pay for a new TLS handshake. And when you interrupt the AI, ElevenLabs drops the request, so the proxy closes the stream and records the turn as cancelled in Sentry.

Here's the proxy at work in Render's logs, one line per Gemma request from the agent:

Render application logs filtered to POST /llm/v1/chat/completions: a steady list of 200 OK responses on October 4 and 5

The service is managed by the Blueprint, which stays synced to the repo, and every push to main deploys on its own:

Render Blueprints page: before-i-call, synced with adkbbx/before_I_call on main

Render events for the before-i-call web service (Docker, Starter, Blueprint managed, live): deploys started by

Fixing how numbers are read

The first version read 7ๆ™‚ as "nana-ji". Nobody says that. Japanese numbers change their reading depending on the counter after them: 7ๆ™‚ is ใ—ใกใ˜, 10ๅˆ† is ใ˜ใ‚…ใฃใทใ‚“, 4ๆœˆ1ๆ—ฅ is ใ—ใŒใคใคใ„ใŸใก and 2ไบบ is ใตใŸใ‚Š. Dates, times and the number of people are exactly what you say on a phone call, so this had to be right.

I fixed it with a pronunciation dictionary: 118 entries I wrote by hand plus 5,989 generated ones for dates, clock times, durations and counters, 6,107 in total. A script builds the rules and uploads them to the ElevenLabs agent as a pronunciation dictionary, and the app uses the same list for slow replay and the recorded demos. The text on screen doesn't change. This is the dictionary on the agent:

ElevenLabs agent settings: the pronunciation dictionary

Japanese voices trip over dates and counters, so I wrote 6,107 fixes: 7ๆ™‚ is shichiji, not nana-ji

Tuning the prompt for Gemma

Switching to Gemma meant rewriting the prompt. I tested it with text-only runs of the agent. The old prompt never called end_call (0 of 2 scenarios), and the tuned one hung up in all 4. It also keeps ordinary words in kanji and writes numbers in kana so the voice reads them correctly, and it asks for practice details instead of your real ones. It still missed the hang-up sometimes in real calls, and Sentry caught that. The start of the tuned prompt:

The start of the agent's system prompt in ElevenLabs: polite spoken Japanese, one question per turn, kanji for ordinary words, numbers with counters in hiragana such as 7ๆ™‚โ†’ใ—ใกใ˜, and never inventing dates, prices or availability

Checking every reply with Sentry

Since every turn goes through my server, I can check each reply against the rules of the practice call. If you said goodbye and the AI didn't hang up, if it "confirmed" a booking, or if it put digits or English letters into Japanese speech, that becomes a Sentry issue tagged with the prompt version and turn number. Each request is also a gen_ai trace with time to first token, token counts and estimated cost. After a call you can tap Report this reply, and the report shows up next to that turn's trace. Conversation text only goes to Sentry if you choose to attach it.

Here's a real one. On October 3, the practice partner didn't hang up after the learner said goodbye, and Sentry caught it 8 times in four minutes. Each rule gets its own issue:

Sentry issue feed showing one unresolved issue, Practice partner: missed hang up, on /llm/v1/chat/completions with 8 events

The issue lists every time it happened and which release it came from:

The missed hang up issue in Sentry: 8 events between 5:57 and 6:01 PM UTC on October 3, all on the same release in production

For comparison, here's a turn where the partner did hang up. The execute_tool end_call span is the hang-up. The Gemma call took 1.27 s (Sentry's average for it is 1.49 s), used about 1,100 input tokens and 43 output tokens, and cost less than a cent. Traces don't include the conversation text, so the Input tab is empty.

A Sentry trace of one practice turn: invoke_agent, chat gemma-4-31B-it taking 1.27 s, the end_call tool and the request to DigitalOcean, with 1.1K input tokens, 43 output tokens and a cost under $0.01

All of those Gemma calls run on DigitalOcean's serverless inference, so I never had to run a GPU server. In two days, October 3 and 4, the app sent about 1.28 million input tokens to Gemma 4 31B and got about 75,000 back. Most of it is input because every turn sends the instructions and the conversation so far, while the replies are short, like the 43 tokens in the trace above.

DigitalOcean Serverless Inference, Analyze tab: 1,278,483 input tokens and 74,719 output tokens, all used on October 3 and 4

Bugs I hit along the way

  • Gemma sometimes wrote voice directions like [calm] or [pause] into its replies. The app strips them from the text on screen.
  • Small models sometimes say a tool call out loud, like end_call(reason="done"), instead of calling the tool. The app removes that text before it's spoken and treats it as a hang-up.
  • Running Gemma 4 locally, the first replies came back empty because the model spent all its tokens on hidden reasoning. Setting reasoning_effort: "none" fixed it, and replies now take about half a second on my laptop's GPU.
  • Kokoro's Japanese voice wouldn't install on Windows, because pyopenjtalk needs CMake to build. pyopenjtalk-plus has prebuilt wheels, so I switched to that.
  • Romaji showed 10ๆœˆ4ๆ—ฅ as "10tsuki4hi". Romaji and furigana now use the same reading list as the voice, so it reads juugatsu yokka.

Why Does Open Innovation Matter?

The hosted version runs on paid APIs. I also wanted a version that's free and private, so the same code can run entirely on a laptop. With LOCAL_VOICE=1, Whisper handles speech recognition, Gemma 4 E2B (4.3 GB, through Ollama) plays the other person, and Kokoro-82M speaks. Explanations, slow replay, Enhance prompt, the word list, furigana and the PDF all still work. Nothing is sent to a provider, it works offline, and you can switch models with one setting (LOCAL_LLM_MODEL).

Same app, two brains: Gemma 4 31B on DigitalOcean with ElevenLabs, or Gemma 4 E2B on a laptop with Whisper and Kokoro

On my laptop (Ryzen 9 5900HS, RTX 3060 with 6 GB), the AI greets you in 2.4 s, answers what you say in about 6 s, and an explanation takes 1.4 s.

A free local practice call: Gemma's reply with furigana and romaji, a tap-to-speak microphone and the call controls

The local version isn't as good. The ElevenLabs voice sounds more natural, responds faster and lets you interrupt it. Locally you tap to talk, and the small model doesn't follow the role-play rules as closely as the 31B. It's possible at all because Gemma's weights are open: the same model family runs on DigitalOcean and on my laptop, with the same prompt, rules and pronunciation list. Furigana and romaji don't need a model. The open-source pykakasi and Janome libraries handle them.

Prize Categories

  • Best Use of ElevenLabs: one agent over WebRTC for every call, with per-call prompt, first line, language and voice overrides, the built-in end_call tool, a custom LLM pointed at Gemma, a separate text-only session for explanations, text to speech for slow replay, and a 6,107-entry pronunciation dictionary on the agent.
  • Best Use of Sentry Agent Tracing: every Gemma request is a gen_ai trace with time to first token, tokens and cost. Replies are checked against the role-play rules, and user reports link to the exact turn.
  • Best Use of DigitalOcean: serverless inference runs Gemma 4 31B for every call turn, explanation, prompt enhancement and word list, with no GPU to manage.
  • Best Use of Gemma: Gemma does four jobs (the other person on the call, the explanation helper, the prompt enhancer and the word picker), as 31B in the cloud and E2B on a laptop.
  • Best Use of Render: one Docker web service from a render.yaml Blueprint serves the app, the API and the custom LLM proxy that streams every Gemma reply to ElevenLabs, with a persistent disk for usage limits and analytics, keys kept out of the repo, and the credit limits as environment variables.

Built solo.

If you live in Japan and put off phone calls too, try it and tell me in the comments what's missing.

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.