AvaClone

Browser SDK

A small TypeScript library that renders the avatar into any element and speaks the realtime protocol. The widget is built on it; use it directly when you want your own interface.

Install

npm install @avaclone/sdk

The package is published with the dashboard; until it is on npm, load the ES module from the CDN:

<script type="module">
  import { AvaClone, MicCapture, CameraCapture } from "https://cdn.avaclone.ai/v1/sdk.js";
</script>

Mount

Your server creates a session with the API and hands the browser its ws_url, token and idle_loop_url. The browser mounts:

import { AvaClone, MicCapture } from "@avaclone/sdk";

const s = await fetch("/my-server/session").then((r) => r.json());   // your server called POST /api/v1/sessions

const avatar = AvaClone.mount(document.querySelector("#stage")!, {
  wsUrl: s.ws_url,
  token: s.token,
  idleLoopUrl: s.idle_loop_url,
  packed: s.packed,          // frames carry alpha (transparent background)
  lazy: true,                // dial only when the visitor begins
});

avatar.on("transcript", ({ role, text }) => console.log(role, text));
avatar.on("speaking", (on) => button.disabled = on);

const mic = new MicCapture();
await mic.start((pcm) => avatar.sendAudio(pcm));   // 24 kHz mono PCM, chunked

avatar.sendText("What are your opening hours?");   // typed questions work too

Mount options

FieldTypeMeaning
wsUrlstringwss://rt.avaclone.ai/v1/sessions/<id>, from POST /sessions.
tokenstringThe session token. Sent as a query parameter on connect; it is swapped server-side and never reaches the engine.
idleLoopUrlstringThe avatar's idle loop (MP4). Shown at once and whenever the stream is not live. Costs no minutes.
packedbooleanThe idle loop carries its alpha in a side-by-side pack and is composited in WebGL. Read from the session when omitted.
lazybooleanDo not connect until the first sendAudio or sendText. Use it for public visitors.
backoffMs[min, max]Reconnect back-off in milliseconds.
audioContextAudioContextShare one AudioContext with the rest of your page (ducking, analysers).
ownerstringA stable id for this viewer so a reconnect resumes the same session.

Methods

FieldTypeMeaning
sendAudio(pcm)Int16Array | ArrayBufferMicrophone audio, 24 kHz, 16-bit mono. The relay transcribes with server-side voice detection; you do not need to detect speech yourself.
endUtterance()Optional: mark the end of what the visitor said when you run your own push-to-talk.
sendText(text)stringA typed turn. Queued until the socket is open.
interrupt()Stop the avatar mid-sentence and clear buffered audio (barge-in). Speaking into the microphone does this automatically.
setCamera(cam | null)CameraCaptureHand the avatar a CameraCapture so the agent can ask for a still when the visitor refers to something in view. null turns it off.
rest()End the session on purpose (the visitor walked away). The idle loop returns.
destroy()Tear everything down and release the element.
on(event, fn)Subscribe to an event below. Returns a function that unsubscribes.

Events

FieldTypeMeaning
ready{ avatar, packed, fps, chunk_s, in_sr, model }The stream is up; the first frame is about to show.
session{ id, kind, agent }The relay accepted the session.
transcript{ role, text, final, via }What the visitor said (role user, via voice or text) and what the agent is saying (role assistant; final is false while it streams).
speakingbooleanThe avatar started or stopped speaking.
frame(seq, kind)A frame was drawn; kind is idle or speech.
stats{ chunks, late, genMsP50, bytesPerS }Playback health, about once a second.
restingThe stream ended (quiet cut-off, rest(), or a cap) and the idle loop is showing.
closed(code, reason)The socket closed. Codes: 4001 quiet cut-off, 4002 the avatar is loading (retry shortly), 4003 a plan cap, 4009 all engines busy, 4401 bad token.
errorErrorSomething failed; the SDK reconnects where it can.

MicCapture and CameraCapture

MicCapture opens the microphone with echo cancellation and noise suppression, resamples to 24 kHz and calls you with PCM chunks. CameraCapture opens the camera (user or environment), optionally previews into an element, and keeps the last three seconds of frames in memory; flip() switches cameras on phones. Neither sends anything by itself: audio goes only through sendAudio, and a camera still goes only when the agent asks for one.

import { CameraCapture } from "@avaclone/sdk";
const cam = new CameraCapture();
await cam.start("environment", document.querySelector("#preview")!);
avatar.setCamera(cam);      // later: avatar.setCamera(null); cam.stop();

Streaming your own speech (stream sessions)

A session of kind stream has no agent: you send 24 kHz PCM of speech you generated (your own assistant, a human operator, a recording) and the avatar lip-syncs to it. Call endUtterance() when a passage ends. Tokens are drawn while the stream is live, 30 a minute on Standard.

Notes

  • The element you mount into keeps its own positioning; the SDK only adds a canvas and video that fill it.
  • Browsers require a user gesture before audio plays. Mount on load, but call sendText or start the microphone from a click.
  • Frames arrive as JPEG at 25 fps with PCM audio in 0.96 s chunks; the SDK schedules both on one clock so lips and sound stay together.