WebSocket calls
A WebSocket call connects a real phone call to a server you run. Audio arrives as JSON messages over a single wss:// connection, and whatever you send back is played to the person on the phone. There is no SIP stack to operate and no media server to deploy — if your code can hold a WebSocket and handle 16-bit audio, it can be on a call.
What you can build
- Voice AI agents on a phone number. Connect your speech pipeline — speech-to-text, an LLM and text-to-speech, or a realtime voice model — straight to calls. Audio is 24 kHz PCM, the format many realtime voice models use, so there is often nothing to resample.
- An AI receptionist. Answer calls to your didlogic number, understand what callers need and respond, with keypresses (DTMF) for menus.
- AI as backup for your team. Put your endpoint after a phone in the number's dialing sequence: your team answers first, and your agent picks up when no one does.
- Outbound calls placed by your agent. Start a call over the API for appointment reminders, delivery confirmations or surveys. Your agent speaks once the person answers.
How it works
It works in both directions:
| Who starts it | Who your endpoint talks to | |
|---|---|---|
| Inbound | somebody dials one of your didlogic numbers | the caller |
| Outbound | you place a call over the Click2Call API | the number you dialed |
Both use the same connection, the same audio format and the same messages. Only the setup and the first few seconds differ, so read Your endpoint and The messages once, then the section for the direction you need.
Before you start
- Ask support to enable WebSocket calls. Inbound and outbound are enabled separately, per account. Until then, inbound WebSocket destinations cannot be saved and outbound WebSocket calls are refused.
- Outbound calls need API access. They are placed over the Click2Call API, which is available from the Plus plan.
Pricing
WebSocket calls cost nothing extra:
- Inbound calls are billed at your number's regular inbound rate.
- Outbound calls are billed as regular calls from the SIP account you place them under.
Your endpoint
Your endpoint is a WebSocket server. The platform connects to it; it never connects to us.
The URL must be **wss://**. Plain ws:// is refused. The certificate must be valid for the hostname — there is no option to skip verification.
The host must be publicly reachable. It has to resolve to a public IPv4 address. These are refused when you save or place the call, not silently at call time:
- loopback and private ranges (
127.0.0.0/8,10/8,172.16/12,192.168/16) - carrier-grade NAT (
100.64/10) and the documentation ranges (192.0.2/24,198.51.100/24,203.0.113/24) localhostand internal-looking names such as.localor.internal- IP literals in the URL in place of a hostname, for IPv6
- hosts that resolve to IPv6 only. A dual-stack host is fine; the connection uses its IPv4 address
Any port is allowed; 443 is assumed when you give none.
You must say **ready** within 5 seconds. As soon as the socket is open you get a start message, and you reply:
{"event": "ready"}
Until that arrives the call is not answered. If you do not send it in time the connection is closed with code 1008 and a reason explaining what was expected, so a silent endpoint shows up in your logs as a refusal rather than as a dropped connection. You do not have to wait for start to send it — what you must not do is wait for anything else first.
This is deliberate: an endpoint that is down, or one that accepts the socket but does not speak this protocol, never produces an answered call of silence that you then pay for.
The audio format
One format, both directions, no negotiation:
| Encoding | PCM, signed 16-bit, little-endian |
| Sample rate | 24000 Hz |
| Channels | 1 (mono) |
| Frame | 10 ms — 240 samples, 480 bytes |
| Transport | base64 of those bytes, in the JSON payload field |
A 10 ms frame is 480 bytes of PCM, which is 640 characters of base64.
What you receive is always exactly one 10 ms frame per media message.
What you send is more forgiving: any whole number of samples is accepted and buffered, so you may send 20 ms or 100 ms chunks, or whatever your text-to-speech hands you. A partial frame is held until the rest arrives. Sending 10 ms frames in real time gives the smoothest result; sending a large burst is fine too, and it plays out at the natural rate.
The messages
Every message is a JSON text frame with an event field and a call_id. Binary WebSocket frames are not used in either direction.
What the platform sends you
| Event | When | Fields |
|---|---|---|
start | immediately after the socket opens | call_id, from, to, codec, sample_rate, frame_bytes, ptime, and on an outbound call direction: "outbound" |
media | continuously, once the call is up | call_id, payload |
dtmf | the person pressed a key | call_id, digit |
answered | outbound only — the destination picked up | call_id |
stop | the call has ended | call_id |
A start looks like this:
{
"event": "start",
"call_id": "01K5ZQJ3M8N7P0R2T4V6W8X9Y0",
"direction": "outbound",
"from": "442071234567",
"to": "447700900000",
"codec": "pcm16",
"sample_rate": 24000,
"frame_bytes": 480,
"ptime": 10
}
Take the audio format from these fields rather than hard-coding it, and ignore any event you do not recognize — new ones may be added, and a client that rejects unknown events will break when they are.
What you send
| Event | Meaning |
|---|---|
ready | you speak this protocol; answer the call. Required |
media | audio to play, in payload |
clear | discard everything you have sent that has not played yet |
hangup | end the call |
{"event": "media", "payload": "AAAAAAAA…"}
clear is what makes interruption work. Audio you send is queued and played out in order, so if the person starts talking over your reply, send clear and the rest of that reply is dropped instead of finishing into the conversation. Then send the new audio.
hangup ends the call from your side. You still get a stop afterwards, as you would for a call the other party ended, and then the socket is closed normally.
Your endpoint cannot send DTMF; dtmf is one-way, towards you.
Inbound calls
Somebody dials one of your didlogic numbers and the call is delivered to your endpoint.
Point a number at your endpoint
A WebSocket endpoint is a destination on one of your numbers, alongside SIP URIs, phone numbers and SIP devices. It is transport 12:
curl -X POST "https://app.didlogic.com/api/v1/purchases/15551230000/destinations" \
-H "Authorization: Bearer ${DIDLOGIC_API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"destination": {
"transport": 12,
"destination": "wss://voice.example.com/didlogic?token=s3cret",
"active": true
}
}'
You can do the same in the customer portal: edit the number's destinations and pick the type WebSocket.
It behaves like any other destination from there on — active and the number's dialing mode (sequence, ring_all, random) work the same way, so you can put your endpoint beside a phone and ring both, or fall back from one to the other.
The secret goes in the URL, as a **token** query parameter. It is taken out of the URL before anything connects and presented as a header instead:
GET /didlogic HTTP/1.1
Authorization: Bearer s3cret
So your server checks the Authorization header. The token parameter is not in the request line — it never reaches your access log, and you can move it later without your endpoint noticing. Every other query parameter you put in the URL is sent as-is.
What happens on a call
- Somebody dials your number.
- Your endpoint gets a connection, with the
Authorizationheader. - You get
start.fromis the caller's number,tois your number that was dialed. - You reply
ready. - The call is answered. The caller is already on the line — they dialed you — so this is the moment to greet them.
mediaflows both ways, anddtmfarrives for any key the caller presses.- The caller hangs up, or you send
hangup, and you getstop.
There is no answered event on an inbound call: nothing is waiting to pick up.
If your endpoint is unreachable, or does not send ready, the call is simply not answered — the caller hears ringing and the call ends unanswered, with nothing billed.
Outbound calls
You place the call. Your endpoint takes one side of it and a phone number takes the other.
This is the Click2Call API with a_type: "stream". Instead of ringing one of your SIP accounts first, the platform connects your endpoint first, and only dials the destination once you are ready.
curl -X POST "https://app.didlogic.com/api/v1/click2call" \
-H "Authorization: Bearer ${DIDLOGIC_API_TOKEN}" \
-H "Content-Type: application/json" \
-d '{
"a_type": "stream",
"a_endpoint": "wss://voice.example.com/didlogic",
"stream_token": "s3cret",
"sip_account": "10001",
"b_number": "+447700900000",
"callback_url": "https://hooks.example.com/didlogic",
"client_call_id": "order-8821"
}'
| Field | Notes |
|---|---|
a_type | "stream" |
a_endpoint | your endpoint's wss:// URL, up to 2048 characters |
stream_token | the secret, presented as Authorization: Bearer. 1–512 printable characters, no spaces. Optional — omit it and your endpoint is connected with no Authorization header. Never returned in any response or webhook |
sip_account | one of your own five-digit SIP accounts. The call is placed and billed as that account, under its limits, and the destination sees its caller ID |
b_number | the destination, E.164; a leading + is stripped |
callback_url, client_call_id and a_cli work exactly as they do on any other Click2Call call. b_cli is still refused.
Here the token does not go in the URL. a_endpoint is dialed exactly as you wrote it, query string included, so a token parameter in it stays an ordinary parameter and is not turned into a header. Put the secret in stream_token. (This is the opposite of an inbound destination, where the URL carries it — the difference is that an outbound request has a field for it.)
The response is the usual call document, with a_type: "stream" and a_endpoint showing the URL you supplied:
{
"data": {
"id": "01K5ZQJ3M8N7P0R2T4V6W8X9Y0",
"state": "accepted",
"a_type": "stream",
"a_endpoint": "wss://voice.example.com/didlogic",
"b_number": "447700900000",
"client_call_id": "order-8821"
}
}
id is the call_id you will see on the socket.
What happens on a call
201 Created. Nothing has rung yet.- Your endpoint gets a connection, with the
Authorizationheader if you sentstream_token. - You get
start, with"direction": "outbound".tois the destination,fromis the caller ID it will see. - You reply
ready. - The destination starts ringing. Not before — your endpoint being up is what releases the call.
- You get
answeredwhen the person picks up. mediaflows both ways until the call ends, thenstop.
Wait for **answered** before you say anything. Audio flows from step 5, not from step 6: before the destination picks up you are hearing ringback, and anything you send goes into a phone that is still ringing. An endpoint that greets on start greets nobody.
If the call ends before the destination answers — busy, no answer, rejected — you get stop with no answered first. The socket does not tell you why; the reason is on the call, in the click2call.* webhooks and in GET /api/v1/click2call/:id.
The platform sends a WebSocket ping every 20 seconds on outbound calls, so a proxy in front of your endpoint does not close an idle-looking socket. Any normal WebSocket library answers these for you.
When the request is refused
All of these are 422, with error.code validation_failed and the specific code in error.details. No call is placed, so nothing is billed and your endpoint is never connected.
| Code | field | Cause |
|---|---|---|
a_type_unsupported | a_type | the account is not enabled for WebSocket calls |
invalid_stream_url | a_endpoint | not wss://, longer than 2048 characters, or a host that fails the rules in Your endpoint |
invalid_stream_token | stream_token | empty, longer than 512 characters, or contains a space or a non-printable character |
unknown_sip_account | sip_account | missing, not five digits, or not one of your own accounts |
invalid_b_number | b_number | not an E.164 number |
They are checked in that order, and a_type_unsupported carries the message a_type must be sip_account — on an account without WebSocket calls enabled, stream is simply not a type that exists.
Ending a call
| What happens | What you see |
|---|---|
| the other party hangs up | stop, then a normal socket close |
you send hangup | stop, then a normal socket close |
| the connection drops | the call ends |
stop is the last message on a call, whoever ended it. Treat the socket as finished once you have it.
A minimal endpoint
This is an echo test: it answers, waits for the call to be live, and plays your own voice back at you. Enough to prove the plumbing before you put a real voice pipeline behind it.
import asyncio
import json
import websockets
async def handler(connection):
# Check the Authorization header here before going any further.
live = False
async for raw in connection:
message = json.loads(raw)
event = message.get("event")
if event == "start":
await connection.send(json.dumps({"event": "ready"}))
# Inbound: the caller is already there, so speak after `ready`.
# Outbound: wait for `answered`.
live = message.get("direction") != "outbound"
elif event == "answered":
live = True
elif event == "media" and live:
await connection.send(json.dumps({
"event": "media",
"payload": message["payload"],
}))
elif event == "dtmf":
print("digit", message["digit"])
elif event == "stop":
break
async def main():
async with websockets.serve(handler, "0.0.0.0", 8080):
await asyncio.Future()
asyncio.run(main())
Terminate TLS in front of this — the platform connects with wss:// only.
Troubleshooting
| What you see | Likely cause |
|---|---|
Saving the destination, or the call, is refused with invalid_stream_url | not wss://, or the host resolves to a private address, an internal-looking name, or IPv6 only |
The socket closes immediately with code 1008 | no {"event":"ready"} within 5 seconds of connecting |
| No connection arrives at all, inbound | the destination is not active, the number has a higher-priority destination answering first, or the account is not enabled |
a_type_unsupported on an outbound call | the account is not enabled for WebSocket calls |
| Calls answer but both sides hear nothing | audio at the wrong rate or width — it must be PCM 16-bit 24000 Hz mono, base64 of raw little-endian bytes, with no WAV header |
| Your greeting is cut off or talks over ringing, outbound | you spoke on start; wait for answered |
| Audio keeps playing after the person interrupts | send clear to drop what has not played yet |
| Nothing arrives for a while and the socket dies | a proxy in front of your endpoint timed out; make sure it passes WebSocket pings through and does not impose its own idle timeout |