Phone calls are the boss fight of living in Japan. ๐ผ
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend I've lived alone in Tokyo for four years now, but when my washing machine started leaking, I stared at the building manager's number for
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
I've lived alone in Tokyo for four years now, but when my washing machine started leaking, I stared at the building manager's number for twenty minutes before I pressed call.
My friend ran into the same thing when she had to cancel her internet contract by phone. The provider's explanation sounded alien to her. The Japanese businesses use on the phone is not the Japanese you use every day, and on a call there's no face to read and no time to look anything up.
So I built Before I Call, an app where you practice the call with an AI before you make the real one. I made it for myself and for friends in the same situation, and sent it to one of them to see if it actually helps. Their reply is further down.
Here it is in action:
What I Built
You describe the call you need to make, or pick one of the examples (home repairs, the clinic, a missed delivery, the city office, dietary requests, lost property, bills). Then you talk to an AI that plays the person on the other end, like a receptionist or your building manager. It speaks polite Japanese, asks one question at a time and waits for your answer. There's no score.
If you get stuck, you don't have to hang up. Explain question opens a helper next to the call with what the question means and a reply you could give. The AI on the call doesn't see it, so the role-play carries on. Slow replay plays the question again at 0.7x speed, and Repeat question asks the AI to say it again. Kanji have furigana, every line has romaji, and known words show their meaning when you hover or tap.
When you say goodbye, the AI hangs up and you get a call card: the whole conversation with readings and meanings, and a list of useful words from your call. You can download it as a PDF and keep it next to you when you make the real call.
It won't make up dates, prices or availability, and it can't book anything for you. It's only for practice.
Setting up a call. You can type or dictate it in English or Japanese, or start from an example. Enhance prompt tidies up rough notes without adding facts you didn't give it.
Asking for help in the middle of the call:
The word list on the call card:
What my friend said
I sent the link to my friend on WhatsApp and asked them to let me know if it helps. This is what came back:
I didn't expect the PDF to be the part they liked. I made it as a cheat sheet for the real call, but for someone studying for the JLPT it's also a word list from a conversation they actually had.
Demo
Try it at before-i-call.onrender.com. There's no sign-up. Voice practice needs a microphone. Without one, tap Play guided example to watch a recorded call in Japanese or English.
Code
Before I Call
Rehearse the phone call you've been putting off.
A patient AI voice partner for everyday Japanese and English calls, with furigana on every word,
help that never interrupts the conversation, and a call card to take with you.
Use it in the cloud, or free and offline on your own computer with Gemma 4, Whisper and Kokoro.
Live app ยท How it works ยท Free local mode ยท Architecture ยท Run it locally
The problem
A phone call is the hardest everyday conversation in a second language. There is no face to read and no time to look anything up, and the other person speaks at native speed. People who live in Japan put off calling the building manager, the clinic or the delivery company, not because they can't manage the conversation, but because the call itself feels risky.
Phrasebooks help with the first sentence. They don'tโฆ
MIT licensed and built during the challenge. It has 87 tests (80 Python, 7 Node), and the README explains how to run it locally.
How I Built It
The call runs on ElevenLabs Agents
The live call is an ElevenLabs agent. It handles all the audio: the WebRTC connection, speech recognition, working out when you've finished speaking, interruptions, and the voice. I use one agent for every scenario. When a call starts, the app sends that call's prompt, first line, language and voice as overrides, so the same agent can be a Japanese building manager or an English clinic receptionist.
I slowed the agent's speech to 0.85x for learners and capped calls at two minutes, with a daily limit on the server so the credits can't all go in one day. Hanging up uses the built-in end_call tool. Its description tells the model to use it only after your question is dealt with and you've said you don't need anything else, not when you just say thanks. Explain question opens a second, text-only session with a helper prompt, so the explanation never ends up in the call transcript. Slow replay uses ElevenLabs text to speech (Flash v2.5 at 0.7x).
The agent has had 113 conversations so far. ElevenLabs groups them by topic on its own, and the top three match the examples in the app: booking an appointment (27), the washing machine (19) and a missed delivery (17).
Why I switched to a custom LLM
When I started, the agent used one of the LLMs hosted by ElevenLabs, Qwen3.5-397B-A17B. Its usage came out of the same ElevenLabs credits I got for Hacktoberfest, and it was going through them fast. ElevenLabs lists it at about $0.016 a minute and $0.003 a message, on top of the voice. A custom LLM costs nothing on the ElevenLabs side:
I wanted those credits for the voice, so I switched the agent to a custom LLM and pointed it at Gemma 4 31B on DigitalOcean, which bills my DigitalOcean account instead. With the per-token prices my tracing code uses for Gemma ($0.18 per million input tokens and $0.50 per million output tokens), the turn in the trace further down cost about $0.0002. The ElevenLabs credits now only pay for the audio.
At first the custom LLM pointed straight at DigitalOcean. Later I put my own server in between so I could trace and check every reply. This is the agent's LLM setting now:
On each turn, the agent sends the conversation to my FastAPI server on Render, the server forwards it to Gemma 4 31B on DigitalOcean's serverless inference, and the response streams straight back, so the voice can start as soon as the first words arrive.
The proxy on Render
The custom LLM URL points at my own server: a FastAPI app that Render runs as one Docker web service. The same service serves the React app, the API and the proxy. The whole setup is a render.yaml Blueprint in the repo (trimmed here):
services:
- type: web
name: before-i-call
runtime: docker
plan: starter
healthCheckPath: /api/health
disk:
name: usage-data
mountPath: /var/data
sizeGB: 1
envVars:
- key: LLM_PROXY_KEY
sync: false
- key: GRADIENT_MODEL
value: gemma-4-31B-it
- key: MAX_CONCURRENT_CALLS
value: '2'
- key: MAX_CALL_SECONDS
value: '120'
- key: MAX_CALLS_PER_VISITOR_DAY
value: '3'
Keys are sync: false, so they're set in the Render dashboard and never committed. The limits that protect my credits are plain environment variables: two calls at once, two minutes each, three calls per visitor a day. The 1 GB disk holds the SQLite files for those limits and the analytics, so they survive redeploys.
The proxy is a single route. It accepts only the agent's key, refuses every model except Gemma 4 31B, and passes DigitalOcean's stream through unchanged while it reads each chunk for tracing and the rule checks (trimmed):
@router.post('/llm/v1/chat/completions')
async def chat_completions(request: Request):
...
if not isinstance(body, dict) or body.get('model') != model:
# The key is never usable for other, more expensive models.
raise HTTPException(400, f'Only {model} is available.')
...
async def relay():
async for chunk in response.aiter_bytes():
stream.feed(chunk)
yield chunk
return StreamingResponse(relay(), media_type='text/event-stream',
headers={'Cache-Control': 'no-cache', 'X-Accel-Buffering': 'no'})
An always-on web service suits this. Each reply is a streamed response that stays open while Gemma talks. The server reuses one HTTPS client for DigitalOcean, so a turn doesn't pay for a new TLS handshake. And when you interrupt the AI, ElevenLabs drops the request, so the proxy closes the stream and records the turn as cancelled in Sentry.
Here's the proxy at work in Render's logs, one line per Gemma request from the agent:
The service is managed by the Blueprint, which stays synced to the repo, and every push to main deploys on its own:
Fixing how numbers are read
The first version read 7ๆ as "nana-ji". Nobody says that. Japanese numbers change their reading depending on the counter after them: 7ๆ is ใใกใ, 10ๅ is ใใ ใฃใทใ, 4ๆ1ๆฅ is ใใใคใคใใใก and 2ไบบ is ใตใใ. Dates, times and the number of people are exactly what you say on a phone call, so this had to be right.
I fixed it with a pronunciation dictionary: 118 entries I wrote by hand plus 5,989 generated ones for dates, clock times, durations and counters, 6,107 in total. A script builds the rules and uploads them to the ElevenLabs agent as a pronunciation dictionary, and the app uses the same list for slow replay and the recorded demos. The text on screen doesn't change. This is the dictionary on the agent:
Tuning the prompt for Gemma
Switching to Gemma meant rewriting the prompt. I tested it with text-only runs of the agent. The old prompt never called end_call (0 of 2 scenarios), and the tuned one hung up in all 4. It also keeps ordinary words in kanji and writes numbers in kana so the voice reads them correctly, and it asks for practice details instead of your real ones. It still missed the hang-up sometimes in real calls, and Sentry caught that. The start of the tuned prompt:
Checking every reply with Sentry
Since every turn goes through my server, I can check each reply against the rules of the practice call. If you said goodbye and the AI didn't hang up, if it "confirmed" a booking, or if it put digits or English letters into Japanese speech, that becomes a Sentry issue tagged with the prompt version and turn number. Each request is also a gen_ai trace with time to first token, token counts and estimated cost. After a call you can tap Report this reply, and the report shows up next to that turn's trace. Conversation text only goes to Sentry if you choose to attach it.
Here's a real one. On October 3, the practice partner didn't hang up after the learner said goodbye, and Sentry caught it 8 times in four minutes. Each rule gets its own issue:
The issue lists every time it happened and which release it came from:
For comparison, here's a turn where the partner did hang up. The execute_tool end_call span is the hang-up. The Gemma call took 1.27 s (Sentry's average for it is 1.49 s), used about 1,100 input tokens and 43 output tokens, and cost less than a cent. Traces don't include the conversation text, so the Input tab is empty.
All of those Gemma calls run on DigitalOcean's serverless inference, so I never had to run a GPU server. In two days, October 3 and 4, the app sent about 1.28 million input tokens to Gemma 4 31B and got about 75,000 back. Most of it is input because every turn sends the instructions and the conversation so far, while the replies are short, like the 43 tokens in the trace above.
Bugs I hit along the way
- Gemma sometimes wrote voice directions like
[calm]or[pause]into its replies. The app strips them from the text on screen. - Small models sometimes say a tool call out loud, like
end_call(reason="done"), instead of calling the tool. The app removes that text before it's spoken and treats it as a hang-up. - Running Gemma 4 locally, the first replies came back empty because the model spent all its tokens on hidden reasoning. Setting
reasoning_effort: "none"fixed it, and replies now take about half a second on my laptop's GPU. - Kokoro's Japanese voice wouldn't install on Windows, because pyopenjtalk needs CMake to build. pyopenjtalk-plus has prebuilt wheels, so I switched to that.
- Romaji showed 10ๆ4ๆฅ as "10tsuki4hi". Romaji and furigana now use the same reading list as the voice, so it reads juugatsu yokka.
Why Does Open Innovation Matter?
The hosted version runs on paid APIs. I also wanted a version that's free and private, so the same code can run entirely on a laptop. With LOCAL_VOICE=1, Whisper handles speech recognition, Gemma 4 E2B (4.3 GB, through Ollama) plays the other person, and Kokoro-82M speaks. Explanations, slow replay, Enhance prompt, the word list, furigana and the PDF all still work. Nothing is sent to a provider, it works offline, and you can switch models with one setting (LOCAL_LLM_MODEL).
On my laptop (Ryzen 9 5900HS, RTX 3060 with 6 GB), the AI greets you in 2.4 s, answers what you say in about 6 s, and an explanation takes 1.4 s.
The local version isn't as good. The ElevenLabs voice sounds more natural, responds faster and lets you interrupt it. Locally you tap to talk, and the small model doesn't follow the role-play rules as closely as the 31B. It's possible at all because Gemma's weights are open: the same model family runs on DigitalOcean and on my laptop, with the same prompt, rules and pronunciation list. Furigana and romaji don't need a model. The open-source pykakasi and Janome libraries handle them.
Prize Categories
-
Best Use of ElevenLabs: one agent over WebRTC for every call, with per-call prompt, first line, language and voice overrides, the built-in
end_calltool, a custom LLM pointed at Gemma, a separate text-only session for explanations, text to speech for slow replay, and a 6,107-entry pronunciation dictionary on the agent. -
Best Use of Sentry Agent Tracing: every Gemma request is a
gen_aitrace with time to first token, tokens and cost. Replies are checked against the role-play rules, and user reports link to the exact turn. - Best Use of DigitalOcean: serverless inference runs Gemma 4 31B for every call turn, explanation, prompt enhancement and word list, with no GPU to manage.
- Best Use of Gemma: Gemma does four jobs (the other person on the call, the explanation helper, the prompt enhancer and the word picker), as 31B in the cloud and E2B on a laptop.
-
Best Use of Render: one Docker web service from a
render.yamlBlueprint serves the app, the API and the custom LLM proxy that streams every Gemma reply to ElevenLabs, with a persistent disk for usage limits and analytics, keys kept out of the repo, and the credit limits as environment variables.
Built solo.
If you live in Japan and put off phone calls too, try it and tell me in the comments what's missing.
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ full credit and traffic to the original publisher.




















