Back to the blog
October 6, 2026·14 min read·Jonas

AI telephony: five days, 68 commits and one night of 40 seconds

How Jarvis learned to make phone calls – and why “the assistant should be able to call” is not a weekend feature.

#Telephony#Behind the scenes#Tech
Tech explained simply: “Can’t ChatGPT do that?”
  • ChatGPT delivers text. A phone call is a live audio stream in both directions. If the answer takes three seconds, the conversation is broken – so you need real-time voice instead of “question – answer”.
  • Someone has to own the line. A real phone number comes from a provider. At Relyper your line plugs into our own phone server (Asterisk, open source) – or you use a voice API.
  • The key must never leave the server. Every call gets a short-lived token so the real API key never reaches users.
  • And then everything around it: who may call, which company owns which line, what a minute costs, who pays – things a chatbot knows nothing about.

It sounded simple at first: the assistant should be able to call and be called. Then you start – and one thing leads to the next. You open a door and ten more are waiting behind it. Each one needed its own fix.

From the outside you see none of it: in the end it is just a phone call that works. That is exactly why I am telling you how much work went into it.

An AI assistant that makes phone calls – sounds like a single API call. In reality it took 51 commits in the main project, 17 more in the phone-server repo, around 5,000 lines of telephony code, a phone server of our own and one night when every call dropped after exactly 40 seconds.

Between October 2 and 6, 2026, Jarvis, the AI assistant of the Relyper AI Factory, learned to make phone calls. This is the whole story, including everything that went wrong.

What Jarvis can do now

  • Make calls: “Call the doctor and ask for an appointment.” Jarvis dials, waits for the “Hello?”, says in the first sentence that he is an AI, does the job and then delivers a summary.
  • Take calls: Jarvis picks up, takes messages, puts favourites through, hangs up on blocked numbers and respects quiet hours.
  • Bots: Bots of a company (a “tenant”) can make calls and send text messages – only with approval from a responsible person.
  • Own lines: Every company brings its own lines and numbers, up to 10 per tenant, with a backup line if one fails.
  • Billing: Every minute is charged in Relyper Coins, either from the personal or the company wallet.

The setup: one call, five stations

Explained simply: A call travels through several systems like a relay race. The phone provider brings it into the phone network, a phone server (Asterisk) takes it, our API controls everything, and a voice model listens and answers. The special part: there is no “speech → text → AI → speech”. The model understands and speaks audio directly. That makes it fast, and you can even interrupt it.

Call chain: phone provider, Asterisk, AI Factory API, OpenAI Realtime; Relyper supplies token and budget to the API

The API is the control centre: it steers Asterisk, routes the audio to the voice model and fetches a token and budget from Relyper before every call.

Layer Technology Job
Phone provider SIP trunk (e.g. sipgate, easybell, fonial) or voice API (Twilio, SignalWire, Telnyx) Numbers, connection to the phone network
Phone server Asterisk 22 in a Docker container (open source, free) Registration with the provider, channels, announcements
Control Asterisk REST Interface (ARI) and WebSocket for the audio Start, answer and end calls
AI Factory API Node/TypeScript Rules, limits, approvals, billing, logging
Voice model OpenAI gpt-realtime, voice marin Listening, speaking, triggering actions (transfer, hang up)
Relyper AI proxy, short-lived tokens, coin wallets Key selection, budget check, billing

Three details that show how much is under the hood:

  • Lines without a restart: The API builds the phone configuration from each company’s line data and loads it while running. That is necessary because Asterisk’s dynamic configuration cannot handle registrations with providers.
  • No AI key in the AI Factory: For every call it fetches a one-time token from Relyper (valid for 120 seconds). Relyper picks the right key and checks membership and balance – if the wallet is empty, there is no token and therefore no call.
  • Backup line: If a line fails, Jarvis automatically redials over the next one.

The first plan fails

There were three attempts at the telephony part:

Approach Result Why
sipgate click-to-call plus SIP stream bridge (Oct 2) dropped The bridge could not start calls on its own; click-to-call rings your own phone first
jambonz (voice gateway) dropped Self-hosting costs money
Asterisk 22 plus any SIP trunk (Oct 3) chosen Open source, runs in a container, works with any provider
Voice API (Twilio, SignalWire, Telnyx) additionally No server of our own needed, can also send text messages

“SIP is a standard” is only true on paper. Every provider wants the caller number differently: sipgate as P-Preferred-Identity with digits only, easybell as P-Asserted-Identity, fonial with +49… in the From field. That is why the code has its own provider profiles. You can find a price comparison of the providers in the telephony guide.

The voice model

  • Chosen: OpenAI gpt-realtime, voice marin. Because the model works with audio directly, there is no detour through text.
  • Configurable per number: model and voice, the assistant’s name (e.g. “Aivra”), its behaviour in plain words and knowledge from text or websites.
  • Separate telephony key: Calls run over a key of their own so their costs show up separately from chat costs.
  • When is the caller done? The default is semantic detection: a model judges whether a thought is complete, so background noise and beeps rarely trigger an answer. The alternative only reacts to volume – faster, but then every noise counts as speech.

Green tests, noisy line

After three days the software was done and tested. The real problems only came with the first real call.

Date Problem Cause Fix
Oct 4 Two settings collided Same variable name for two purposes No AI key in the AI Factory anymore, one-time tokens instead
Oct 5 Loud noise in the announcement Playback via ARI was noisy; Git had treated the audio files as text Announcement via the dialplan, audio in raw format, binary files in .gitattributes
Oct 5 New announcements never arrived Docker kept old files in the volume Files are copied fresh on every start, build without cache
Oct 5 API could not find Asterisk Two different Docker networks Asterisk moved into the project network, reached by container name
Oct 5 Scanners flooded the log The internet probes open SIP ports Only provider addresses may connect
Oct 5 sipgate rejected calls sipgate ignores From and PAI Caller as P-Preferred-Identity, digits only
Oct 6, 2 am Calls dropped after about 40 seconds One huge audio frame, see below At most 10 frames per message
Oct 6, 3 am Voice model dropped out mid-call Long sessions, expired tokens Up to 2 reconnects with the conversation so far
Oct 6, 10 am Calls ended for no reason Harmless “error” messages from the model closed the connection Errors are only logged now; only a real disconnect counts
Oct 6 Jarvis answered background noise Detection based on volume only Semantic detection plus a prompt rule “ignore side conversations”
Oct 6 Goodbye got cut off The model hung up while still talking 6 seconds of run-out time

The 40-second night

Everything worked, but every call died after about 40 seconds. The cause was neither the network nor the model: when Asterisk briefly paused and then resumed the audio, the entire backlog went out as one single huge WebSocket frame – and Asterisk closed the connection.

// Asterisk closes the connection when a WebSocket frame gets too big
const MAX_FRAMES_PER_MESSAGE = 10;

I only found it because I had built in diagnostics beforehand: close code, bytes sent, queue size, pause counter. Without them I would have kept guessing for a long time.

One single oversized chunk of data – and every call was dead after 40 seconds.

“Loud noise on the line”

The announcement “The assistant is not available right now” sounded like a radio stuck between two stations. Two causes at once: Git’s line-ending conversion had damaged the audio files, and playback via ARI added noise on top.

The file that would not update

New image built, old announcement still playing. Docker held on to the old files in the volume. Six commits for a single problem followed: first copying every file by name instead of whole folders, plus a build check that fails on an outdated state. Then moving the files out of the Docker volumes entirely, so the start script copies them into place fresh on every start – plus a safety net in case an old template ever slips through. And when the announcement finally arrived, it was too quiet – so it had to get louder.

Six attempts for one announcement – and in the end it was simply too quiet.

The internet knocks

As soon as port 5060 was open, scanners showed up trying to guess passwords for SIP accounts. A public phone server gets attacked immediately, so the access list was mandatory from day one.

The second night: the sentence that never arrived

A day later, silence again – this time in the middle of a paragraph. The transcript had the whole sentence, the phone only the first half. So the model had finished speaking; the audio got stuck on its way to Asterisk.

The cause was, of all things, the fix from the first night. The voice model produces speech faster than it is spoken. I now sent the backlog in small pieces, but as fast as possible. Asterisk then says “buffer full, wait” (MEDIA_XOFF), and if the “go on” (MEDIA_XON) gets lost on the way, the queue waits forever. No error, no message – just silence.

Two measures: audio now goes out paced to real time, at most about a second ahead, so the buffer never fills up. And a watchdog: if no “go on” arrives for a second and a half after a pause, we send anyway. Better a short click than a mute assistant.

When the model loses the connection

If the connection to the voice model really drops, the call used to simply end. Now the AI reconnects, gets the conversation so far, tells the caller the line was briefly interrupted – and carries on exactly where it stopped. Only when that fails twice does the caller hear an announcement before the line is closed. Never a dead line.

I also learned that “error” does not mean “broken”: the model regularly reports harmless errors, for instance when the greeting and an automatic reply collide. For one version I treated every such error as a disconnect and rebuilt the connection. The result: Aivra said “hello” and went quiet. Since then only a real connection loss triggers the reconnect; errors are logged.

The note that did not exist

If the caller hung up before the AI could sum up, the monitor only showed “The caller hung up”. That an appointment had just been arranged was gone – although the full transcript was there. Now every call shows the summary on top and the transcript below. And if the AI did not get to it, a text model writes the note afterwards from the transcript: who, why, every question with its answer, every appointment with date and time.

The key that never paid

Companies store their own OpenAI key for calls at Relyper and tick the models it may pay for. One company had ticked gpt-realtime-2.1 – but the AI Factory asked for gpt-realtime, hardcoded. The comparison is exact, the key was skipped, and every call quietly ran on Relyper’s key.

Two mistakes at once: the model must not be hardcoded in the software, and a fallback must never happen silently. Now the model comes from the list Relyper reports for the paying key – selectable per line and per number, but never outside that list. If a model does not match the key, there is a clear error instead of a silent detour. And which key paid is shown with every call in the monitor: company key or Relyper key.

One more side effect of the model switch: German with an American accent. The model had to be told to speak standard German like a native speaker, and the transcription was given the language – before that it occasionally took a German sentence for Dutch.

Making it feel like a conversation

Once the call works technically, the fine-tuning starts. A hard cut in the middle of a sentence feels broken, so 30 seconds before the soft time limit (5 minutes) there is a gentle hint, and the AI wraps up the conversation itself. The hard limit is 10 minutes. And the goodbye no longer gets cut off, because the line is only closed after a short run-out time.

How we test

Level What Result
Unit tests Fake voice model, real WebSockets, fake Asterisk; rules, limits, failover, billing 58 tests green
End-to-end Real Asterisk plus a second Asterisk as a fake provider, two companies with their own accounts 16 of 16 checks green
Company wallet Real database, parallel payouts 13 checks green
Real calls Line test, test call, AI conversations via sipgate Found the noise, the 40-second drop, the token drop, the caller-ID rejection, the mute half-sentence, the wrong key

The product also has its own test tools: the line test without AI (the line calls you, plays an announcement and then an echo – you hear yourself, which checks registration, outgoing call and audio in both directions in one go), a test call that deliberately starts with “This is Relyper” so the person called does not hang up straight away, and an AI test call with a choice of model and line. Every finished call stores why it ended.

Coins and costs

Every started minute costs 5 cents in Relyper Coins by default, and before the call the payer has to cover at least one minute. Each company can choose who pays: the person who starts the call or the company wallet. The charge is booked after the call, and never twice. Line tests, test calls and announcements when the model is unavailable are free. Limits guard against runaway costs: 5 AI calls per hour, the time limits above and only allowed country codes (default +49).

The result

On October 6, Jarvis holds stable conversations of up to 10 minutes, survives dropouts of the voice model, ignores background noise and says goodbye without being cut off. For personal calls the monitor only stores time, numbers, person and duration – no transcripts. The AI notice always comes in the first sentence.

What I learned:

  1. Green tests do not mean the phone sounds good. 16 of 16 checks passed, but noise, frame sizes and provider quirks only show up in a real call.
  2. Audio is binary. Git, Docker volumes and build caches do not treat files the way you think.
  3. Every provider is different. A standard on paper, a special case in practice.
  4. An error is not a disconnect. Many “error” messages from the model are harmless.
  5. A public SIP port gets attacked immediately. Access list from the start.
  6. Conversations need soft limits. A hint before the end works better than a hard cut.
  7. Build diagnostics before you need them. Only close code, bytes and queue size in the log narrowed down the 40-second bug – and the reason for every call ending has been shown in the monitor ever since, not only in the server log.
  8. Yesterday’s fix is tomorrow’s bug. Small frames instead of one big one: right. All of them at once instead of paced: the next bug.
  9. Never fall back silently. A key that does not match must raise an error – not quietly reroute to another one.
  10. Nothing that was said may get lost. The transcript is there; whoever hangs up still deserves a note.

What comes next

Testing easybell and fonial with real accounts, real accounts at Twilio, SignalWire and Telnyx, an end-to-end run for multiple lines and WhatsApp Business. And Jarvis has a face now: Aivra.