Browser SDK
A small TypeScript library that renders the avatar into any element and speaks the realtime protocol. The widget is built on it; use it directly when you want your own interface.
Install
npm install @avaclone/sdkThe package is published with the dashboard; until it is on npm, load the ES module from the CDN:
<script type="module">
import { AvaClone, MicCapture, CameraCapture } from "https://cdn.avaclone.ai/v1/sdk.js";
</script>Mount
Your server creates a session with the API and hands the browser its ws_url, token and idle_loop_url. The browser mounts:
import { AvaClone, MicCapture } from "@avaclone/sdk";
const s = await fetch("/my-server/session").then((r) => r.json()); // your server called POST /api/v1/sessions
const avatar = AvaClone.mount(document.querySelector("#stage")!, {
wsUrl: s.ws_url,
token: s.token,
idleLoopUrl: s.idle_loop_url,
packed: s.packed, // frames carry alpha (transparent background)
lazy: true, // dial only when the visitor begins
});
avatar.on("transcript", ({ role, text }) => console.log(role, text));
avatar.on("speaking", (on) => button.disabled = on);
const mic = new MicCapture();
await mic.start((pcm) => avatar.sendAudio(pcm)); // 24 kHz mono PCM, chunked
avatar.sendText("What are your opening hours?"); // typed questions work tooMount options
| Field | Type | Meaning |
|---|---|---|
| wsUrl | string | wss://rt.avaclone.ai/v1/sessions/<id>, from POST /sessions. |
| token | string | The session token. Sent as a query parameter on connect; it is swapped server-side and never reaches the engine. |
| idleLoopUrl | string | The avatar's idle loop (MP4). Shown at once and whenever the stream is not live. Costs no minutes. |
| packed | boolean | The idle loop carries its alpha in a side-by-side pack and is composited in WebGL. Read from the session when omitted. |
| lazy | boolean | Do not connect until the first sendAudio or sendText. Use it for public visitors. |
| backoffMs | [min, max] | Reconnect back-off in milliseconds. |
| audioContext | AudioContext | Share one AudioContext with the rest of your page (ducking, analysers). |
| owner | string | A stable id for this viewer so a reconnect resumes the same session. |
Methods
| Field | Type | Meaning |
|---|---|---|
| sendAudio(pcm) | Int16Array | ArrayBuffer | Microphone audio, 24 kHz, 16-bit mono. The relay transcribes with server-side voice detection; you do not need to detect speech yourself. |
| endUtterance() | Optional: mark the end of what the visitor said when you run your own push-to-talk. | |
| sendText(text) | string | A typed turn. Queued until the socket is open. |
| interrupt() | Stop the avatar mid-sentence and clear buffered audio (barge-in). Speaking into the microphone does this automatically. | |
| setCamera(cam | null) | CameraCapture | Hand the avatar a CameraCapture so the agent can ask for a still when the visitor refers to something in view. null turns it off. |
| rest() | End the session on purpose (the visitor walked away). The idle loop returns. | |
| destroy() | Tear everything down and release the element. | |
| on(event, fn) | Subscribe to an event below. Returns a function that unsubscribes. |
Events
| Field | Type | Meaning |
|---|---|---|
| ready | { avatar, packed, fps, chunk_s, in_sr, model } | The stream is up; the first frame is about to show. |
| session | { id, kind, agent } | The relay accepted the session. |
| transcript | { role, text, final, via } | What the visitor said (role user, via voice or text) and what the agent is saying (role assistant; final is false while it streams). |
| speaking | boolean | The avatar started or stopped speaking. |
| frame | (seq, kind) | A frame was drawn; kind is idle or speech. |
| stats | { chunks, late, genMsP50, bytesPerS } | Playback health, about once a second. |
| resting | The stream ended (quiet cut-off, rest(), or a cap) and the idle loop is showing. | |
| closed | (code, reason) | The socket closed. Codes: 4001 quiet cut-off, 4002 the avatar is loading (retry shortly), 4003 a plan cap, 4009 all engines busy, 4401 bad token. |
| error | Error | Something failed; the SDK reconnects where it can. |
MicCapture and CameraCapture
MicCapture opens the microphone with echo cancellation and noise suppression, resamples to 24 kHz and calls you with PCM chunks. CameraCapture opens the camera (user or environment), optionally previews into an element, and keeps the last three seconds of frames in memory; flip() switches cameras on phones. Neither sends anything by itself: audio goes only through sendAudio, and a camera still goes only when the agent asks for one.
import { CameraCapture } from "@avaclone/sdk";
const cam = new CameraCapture();
await cam.start("environment", document.querySelector("#preview")!);
avatar.setCamera(cam); // later: avatar.setCamera(null); cam.stop();Streaming your own speech (stream sessions)
A session of kind stream has no agent: you send 24 kHz PCM of speech you generated (your own assistant, a human operator, a recording) and the avatar lip-syncs to it. Call endUtterance() when a passage ends. Tokens are drawn while the stream is live, 30 a minute on Standard.
Notes
- The element you mount into keeps its own positioning; the SDK only adds a canvas and video that fill it.
- Browsers require a user gesture before audio plays. Mount on load, but call
sendTextor start the microphone from a click. - Frames arrive as JPEG at 25 fps with PCM audio in 0.96 s chunks; the SDK schedules both on one clock so lips and sound stay together.