Docs.

Two endpoints. POST /v1/decide: text in, decision out. WS /v1/listen: stream voice, get decisions while they talk. Everything below is what the server does, checked against the code.

Quickstart

Get a key at megachadcua.com/app. New accounts get $2 free, no card. Then one call:

# export MEGACHAD_API_KEY=mcc_live_…
curl https://api.megachadcua.com/v1/decide \
  -H "Authorization: Bearer $MEGACHAD_API_KEY" \
  -H "content-type: application/json" \
  -d '{"schema": {"action": {"question": "What should the agent do?",
       "options": ["Click Sign in", "Click Pricing", "Scroll down", "Nothing yet"]}},
       "text": "take me to sign in"}'

# {"session_id": "dcd_…",
#  "fields": {"action": {"value": "Click Sign in", "confidence": 0.98, "probabilities": {…}}},
#  "transcript": "take me to sign in", "latency_ms": 112, "model_ms": 15.2, "usage": {"credits": 1}}

Voice works the same way: send the same schema in a start message over a WebSocket, stream the mic, read fields events. See Listen, or skip to megachad.listen_mic.

Base URL: https://api.megachadcua.com. Every request and response body is JSON. CORS is open on /v1/*.

Authentication

Keys look like mcc_live_…. Get, roll and revoke them at /app (up to 10 live keys). Send the key one of three ways:

WhereHowUse it for
HeaderAuthorization: Bearer mcc_live_…Servers. REST and WebSocket.
Headerx-api-key: mcc_live_…Clients that can't set Authorization.
Subprotocolnew WebSocket(url, ["megachad", "mcc_live_…"])Browsers. They can't set headers on a WebSocket.

?key= in the URL is not accepted. URLs end up in logs. You get invalid_key (401 / close 4001).

Rolling a key kills the old one at once, including open sockets: they get invalid_key and close 4001 within a minute.

A key in a public web page is a key anyone can spend. Use the subprotocol for internal tools and prototypes. For public sites, keep the key on your server and proxy /v1/listen, or call /v1/decide from your backend.

Decide (text)

POST /v1/decide. One decision from text. One model pass: about 15 ms on the GPU, 100 to 250 ms round trip.

Request

FieldTypeWhat it does
schemaobject, requiredThe decisions to make. See Schema. Max 64 KB as JSON.
textstring or list, requiredWhat was said. A string, or a conversation: [{"speaker": "user", "text": "…"}, …]. Max 200,000 characters.
speakerstring"user" (default) or "other", when text is a string.
timeout_msintegerHow long to wait for the model. Default 15,000. Clamped to 1,000 to 60,000.

Full example

Request
curl https://api.megachadcua.com/v1/decide \
  -H "Authorization: Bearer $MEGACHAD_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "schema": {
      "action":  {"question": "What should the agent do?", "options": ["Click Sign in", "Click Pricing", "Scroll down", "Nothing yet"]},
      "confirm": {"question": "Did the user confirm?", "type": "yes_no"},
      "urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]}
    },
    "text": "yes, sign me in, quick"
  }'
200 response
{
  "session_id": "dcd_8Hq2vT0bXk1mPz4R",
  "fields": {
    "action": {
      "value": "Click Sign in",
      "confidence": 0.98,
      "probabilities": {"Click Sign in": 0.98, "Click Pricing": 0.01, "Scroll down": 0.006, "Nothing yet": 0.004}
    },
    "confirm": {"value": true, "confidence": 0.93},
    "urgency": {"value": "high", "level": 2, "confidence": 0.81}
  },
  "transcript": "yes, sign me in, quick",
  "latency_ms": 112,
  "model_ms": 15.2,
  "usage": {"credits": 1}
}

Response

  • fields: one entry per schema field. Always value and confidence (0 to 1). Options add probabilities per option; scales add level. See field kinds.
  • usage.credits: what this call cost. 1 for normal calls, more for big ones.
  • latency_ms: our end, request in to answer out. model_ms: GPU time.
  • session_id: dcd_…. Quote it when you email us.

If the fast path is down, the call falls back to a short voice-style session (~1.5 s). Same response shape, but confidence may be null and there are no probabilities. Rare.

Conversations

Two speakers
{"schema": {"agreed": {"question": "Did they agree on Tuesday?", "type": "yes_no"}},
 "text": [{"speaker": "user",  "text": "can we do Tuesday?"},
          {"speaker": "other", "text": "yes, Tuesday works"}]}

Errors come back as {"error": {"code": "…", "message": "…"}} with an HTTP status. Failed calls are free. Full list in Errors.

Listen (voice)

wss://api.megachadcua.com/v1/listen. Stream audio (or text) in, get fields events the moment a decision changes. Usually mid-sentence.

The sequence

  1. YouOpen the socket with your key (header or subprotocol).
  2. YouSend start right away, within 5 s. It carries the schema and audio format.
  3. Serverqueued, only if every model slot is busy. Repeats as you move up. Waiting in line.
  4. Serverready with your session_id. Billing starts here. Must come within 20 s or the socket closes, unbilled.
  5. YouStream binary PCM16 audio in real time. Send text and speaker messages whenever.
  6. Serverfields whenever a decision changes, transcript as words land.
  7. YouSend end when there's no more input.
  8. Serverdone: final value for every field plus the full transcript.
  9. Serverusage: seconds and credits for the session.
  10. ServerCloses with 1000. Anything else is an error close, preceded by an error event.

If the key, plan or limits say no, the socket still opens, then you get one error event and the close code (4001, 4002, 4029 or 1013). Nothing is billed.

Client messages

start first message, exactly once
start
{
  "type": "start",
  "schema": {
    "action":  {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Go back", "Nothing yet"]},
    "confirm": {"question": "Did the user confirm?", "type": "yes_no"},
    "urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]}
  },
  "audio": {"encoding": "pcm_s16le", "sample_rate": 16000},
  "transcript": true,
  "min_change": 0.05,
  "speaker": "user"
}

audio.sample_rate: 8,000 to 48,000; 16,000 is best. transcript: false turns off transcript events. min_change: only report a field when its confidence moves this much. speaker: who talks first. A second start is ignored with a bad_message error: the schema is fixed for the session.

Audio binary frames

Raw 16-bit little-endian mono PCM at the sample rate you declared. Any frame size up to 256 KB. Send it as it's recorded: the server allows 1.5x real time after a 5 s burst, then closes with too_fast. Silence counts as frames, so a live mic never idles out.

text
text
{"type": "text", "text": "click sign in"}

Words as if heard. For typed input, or your own speech recognition.

speaker
speaker
{"type": "speaker", "speaker": "other"}

Who is talking from now on: "user" or "other". See Speakers.

end
end
{"type": "end"}

No more input. The server answers done, then usage, then closes. Closing the socket yourself works too; you just don't get done.

Server messages

queued
{"type": "queued", "position": 2, "eta_seconds": 30, "message": "Every slot is busy. You're #2 in line."}
ready
{"type": "ready", "session_id": "lsn_Q7c1mZp0Lr8sK2aT"}

The model is listening. Billing starts.

fields
{"type": "fields", "fields": {"action": {"value": "Scroll down", "confidence": 0.97}}, "latency_ms": 15.1}

Only the fields that changed. Act on it when confidence is high enough for you; 0.85 to 0.9 is a good start.

transcript
{"type": "transcript", "final": "scroll down", "partial": "a bit more"}

final is settled text; partial is words still being recognized.

done
{"type": "done", "fields": {"action": {"value": "Scroll down", "confidence": 0.98}, "confirm": {"value": true, "confidence": 0.95}, "urgency": {"value": "high", "confidence": 0.83}}, "transcript": "scroll down, yes, it's urgent"}

After end: every field, plus the full transcript.

usage
{"type": "usage", "session_id": "lsn_Q7c1mZp0Lr8sK2aT", "seconds": 12, "credits": 120}

Right before the socket closes, however it closes.

error
{"type": "error", "code": "too_fast", "message": "Audio is arriving faster than real time for 16000 Hz. …"}

Fatal ones are followed by a close. bad_message isn't fatal: that one message was dropped. Codes in Errors.

A plain GET /v1/listen (no upgrade) returns this protocol as JSON.

Schema & field kinds

A schema is an object. Keys are your field names; each value is one of three kinds. Same schema for /v1/decide and /v1/listen.

All three kinds
{
  "action":  {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]},
  "confirm": {"question": "Did the user confirm?", "type": "yes_no"},
  "urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]}
}
KindDeclareValueExtras
Options"options": [...]The option text, exactly as you sent it.probabilities: option text to probability (/v1/decide).
Yes / no"type": "yes_no"true or false. On sockets, null until the conversation touches it, unless you add "unknown": false. On /v1/decide, always true or false.
Scale"scale": [...], low to highThe level text.level: its 0-based index.
  • question is what the model answers. Optional (defaults to the field name). Write one anyway.
  • Options and scales can't be empty. A field with none of the three is rejected (400 bad_schema on /v1/decide).
  • Give option lists an out like "Nothing yet", so the model isn't forced to pick an action while the user is still talking.
  • Max 64 KB of schema JSON.
  • The schema is fixed for a socket. New screen, new options: open a new session.

Speakers

Two speakers: "user" (the default, the person driving the agent) and "other" (anyone else on the call). Anything else is rejected: bad_message on sockets (the message is dropped), 400 bad_request on /v1/decide.

  • Sockets: set "speaker" in start, or send {"type": "speaker", "speaker": "other"} any time. It applies to everything after it.
  • /v1/decide: "speaker" next to a string text, or a list of {"speaker", "text"} lines.
  • Transcripts: once other has been set, transcript lines come back prefixed user: and other: . Single-speaker sessions get no labels.

Waiting in line

Voice sessions stream, so the model has a fixed number of voice slots (see /v1/status). Text decisions take ~100 ms each and have their own, bigger pool, so open voice sessions never block /v1/decide. When a pool is full:

  • Voice waits. Your socket stays open and you get {"type": "queued", "position": 2, "eta_seconds": 30} every time your position changes. First come, first served. When a slot frees up, the start you already sent is replayed and ready follows. Audio and text sent while waiting are dropped, and waiting is never billed. After 2 minutes in line: at_capacity, close 1013.
  • Text doesn't queue. When the text pool is full, /v1/decide returns 503 at_capacity with Retry-After: 2, after at most ~2 s of retrying on our side. Not billed. Retry it.
  • The 5 s start and 20 s ready clocks start when you get a slot, not when you connect.
  • Free accounts never take the last free slot. Add a card and you can.

Limits

WhatLimitOver it
startWithin 5 s of getting a slot. Nothing may come before it.no_start, 1008
readyWithin 20 s of getting a slotno_ready, 1013, not billed
Idle socket60 s with no frames at allidle, 1000
Session length2 hours from readymax_duration, 1000
Audio rate1.5x real time of the declared sample_rate, after a 5 s bursttoo_fast, 1008
Audio frame256 KBbad_message, frame dropped
Sent while the model connects2 MBtoo_fast, 1008
Text frame64 KBbad_message, frame dropped
Text frame rateBurst of 200, then 50 a secondtoo_fast, 1008
Schema64 KB of JSONbad_schema (socket), 413 too_large (decide)
Decide text200,000 characters413 too_large
Decide timeouttimeout_ms 1,000 to 60,000 (default 15,000)Clamped
Open sessionsVoice: free credits 1, with a card 2. Text decisions in flight: free 2, with a card 8. Counted separately.concurrency_limit, 429 / 4029
New voice sessionsFree credits: 30 per 10 minutesslow_down, 429 / 4029
Waiting in line2 minutesat_capacity, 1013
API keys10 live keys per accountRevoke one at /app

Need more than 2 at a time? Email hello@megachadcua.com.

Errors & close codes

REST errors: {"error": {"code": "…", "message": "…"}} with the HTTP status. Socket errors: an {"type": "error", "code", "message"} event, then the close code. Switch on code; messages can change.

Billed? A failed request is never billed. On a socket, time after ready is billed up to the close, whatever closes it. "Time served" below means exactly that.

CodeHTTP / closeMeaningBilled?What to do
invalid_key401 / 4001Key missing, wrong, revoked, or sent as ?key=. Mid-session: the key was rolled.No / time servedUse a live key from /app, in a header or the subprotocol.
payment_required402 / 4002Free credits used up and no card. Open sockets get it at the next minute's check.No / time servedAdd a card at /app.
payment_failed402 / 4002Your last payment failed.NoUpdate the card at /app.
spend_limit402 / 4002You hit the monthly limit you set. Open sockets get it at the next minute's check.No / time servedRaise or remove it at /app, or wait for the 1st (UTC).
concurrency_limit429 / 4029Too many sessions open on your account.NoClose one, or wait for it to finish.
slow_down429 / 4029Free account opened 30+ voice sessions in 10 minutes.NoWait, or add a card.
at_capacity503 / 1013Every slot is busy. Text: within ~2 s. Voice: after 2 min in line.NoRetry after Retry-After (2 s), or reconnect.
model_unavailable503 / 1013Model backend unreachable.NoRetry in a few seconds. Check status.
model_offline503 / 1013Model backend is switched off.NoRetry later. Check status.
no_ready1013The model didn't start within 20 s.NoReconnect.
bad_request400Body isn't JSON, no schema, no text, or an unknown speaker. Also when the model rejects the input.NoFix the body.
bad_schema400 / 1008Decide: a field without options, scale or "type": "yes_no", or an empty list. Socket: a schema over 64 KB.NoFix the schema.
too_large413Schema over 64 KB or text over 200,000 characters.NoSplit it.
no_start1008No start within 5 s, or audio or text before it.NoSend start first, right after connecting.
too_fast1008Audio over 1.5x real time, more than 50 text frames a second, or over 2 MB sent before ready.Time servedStream as recorded. Set audio.sample_rate right. Batch text.
bad_messagesocket stays openOne message dropped: not JSON, text frame over 64 KB, audio frame over 256 KB, unknown speaker, or a second start.n/aFix that message. The session goes on.
idle1000No frames at all for 60 s.Time servedKeep the mic streaming (silence counts), or reconnect when there's input.
max_duration10002 hours since ready.Time servedOpen a new session.
session_lost1011The session expired on our side.Time servedOpen a new session.
model error502 / 1011The model failed or didn't answer in time. Decide codes: timeout, model_closed, model_error.No / time servedRetry.
not_found404No such endpoint.NoEndpoints: POST /v1/decide, WS /v1/listen, GET /v1/status.

Close codes at a glance

CloseMeansRetry?
1000Normal: after done, or idle / max_durationOpen a new one when you need it
1008Your input broke a rule: bad_schema, no_start, too_fastNot until you fix it
1011Model or gateway failed after ready. Time served is billed.Yes
1013Busy or unavailable: at_capacity, model_unavailable, no_ready. Not billed.Yes, in a few seconds
4001invalid_keyWith a valid key
4002payment_required / payment_failed / spend_limitAfter adding a card or raising the limit
4029concurrency_limit / slow_downAfter closing a session, or waiting

Pricing math

Everything is credits. 1 credit = $0.0001. One meter, one line on a monthly Stripe invoice.

  • Voice: 10 credits a second = 6¢ a minute. Wall-clock time from ready to close, talking or not (the session holds a GPU slot). Per second, rounded up, 1 s minimum.
  • Text: 1 credit per decision = 10¢ per 1,000. A call over 4,000 input tokens counts as ceil(tokens / 4000) decisions. Input tokens = ceil(UTF-8 bytes / 4) of the text plus the same for the schema JSON.
  • Free: failed calls, sessions that never reach ready, time waiting in line.
  • New accounts: $2 free = 20,000 credits = about 33 minutes of voice, or 20,000 decisions. Spent first. After that, add a card or calls get payment_required.
  • Spend limit: optional monthly cap in the console (Billing). We email at 80% and at 100%; at 100% calls get spend_limit until the 1st (UTC). No limit unless you set one.
ExampleMathCreditsCost
90-second voice session90 s × 109009¢
90.2-second voice sessionrounds up to 91 s × 109109.1¢
1,000 decide calls, normal size1,000 × 11,00010¢
One 10,000-token decide callceil(10,000 / 4,000)30.03¢
Socket that never got readynot billed00¢
2 minutes in line, then at_capacitynot billed00¢
An hour-long voice session3,600 s × 1036,000$3.60

Every /v1/decide response carries usage.credits. Every socket ends with a usage event. Your console shows the running month.

Status endpoint

GET https://api.megachadcua.com/v1/status. No key. Cached for 10 s.

200 response
{
  "model": "up",
  "slots": 3,
  "in_use": 1,
  "free": 2,
  "demo": true,
  "today": {"decisions": 1834, "fastest_ms": 14.2}
}
  • model: "up", "down" or "unconfigured".
  • slots, in_use, free: voice slots in total, taken, open. free: 0 means new voice sessions wait in line. Text decisions don't use these.
  • today: calls served since midnight UTC and the fastest model time.

History

GET https://api.megachadcua.com/v1/status/history?days=90. No key. Cached for 60 s. days is 1 to 90, default 90. This is what /status draws.

200 response
{
  "days": [{"date": "2026-10-07", "up_pct": 99.65, "checks": 288}],
  "last24h": [{"at": 1791374700000, "model": "up", "decide_p50_ms": 96}],
  "incidents": [{"start": 1791331200000, "end": 1791332700000, "minutes": 25}]
}
  • We check every 5 minutes: the same probe as /v1/status, plus the median wall time of /v1/decide calls in those 5 minutes.
  • days: oldest first, UTC days, today last. up_pct is the share of checks that found the model up, null on a day with no checks.
  • last24h: every check in the last 24 hours. decide_p50_ms is null when no decisions finished in that window.
  • incidents: the latest 10 runs of failed checks in the window, newest first. Times are ms since epoch. end is the first check after that found the model up, null while it's still down.

SDKs & example

Both SDKs pace audio to real time, map every error code to a typed exception with a retry hint, and read MEGACHAD_API_KEY from the environment.

Python

pip install "megachad[mic]"
import megachad   # Python 3.9+. Linux mic: apt install libportaudio2
schema = {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Go back"]}}

# text
d = megachad.decide(schema, "take me to sign in")
print(d["action"].value, d.credits)   # Click Sign in 1

# voice: talk, get decisions as you go
for f in megachad.listen_mic(schema):
    print(f)   # action='Go back' (0.97)

pip install megachad without [mic] for text and raw audio only. async with megachad.listen(schema) as s gives you the socket: send_audio, say, speaker, end. CLI: megachad decide --text "take me to pricing" --options "Pricing,Docs,Nothing yet".

JavaScript

npm i megachad
// Browsers, Deno, Bun, Node 22+. Node 18/20: npm i megachad ws
import { decide, listen } from "megachad";
const schema = { action: { question: "What should the agent do?", options: ["Click Sign in", "Scroll down", "Go back"] } };

const d = await decide({ schema, text: "take me to sign in" });
d.fields.action.value;          // "Click Sign in"

const s = listen({ schema, onFields: (e) => console.log(e.fields) });
await s.ready;                  // throws InvalidKey, PaymentRequired, AtCapacity…
s.sendAudio(pcm16);             // 16 kHz mono PCM16, paced for you
const done = await s.end();     // final fields + transcript

Browser mic in one call: await listenMic({ key, schema, onFields }). It authenticates with the subprotocol, never the URL. Mind the key-in-a-web-page warning.

Example: voice-browser

Drive a real browser with your voice in ~60 lines of Python. It reads the links and buttons on screen, sends them as options, and clicks the one you mean once confidence hits 0.85.

examples/voice-browser
pip install -r requirements.txt && playwright install chromium
export MEGACHAD_API_KEY=mcc_live_…
python voice_browser.py https://news.ycombinator.com

# no mic: typed commands over /v1/decide
printf 'open the top story\ngo back\nscroll down\n' | python voice_browser.py https://news.ycombinator.com --text

Say "open the top story", "go back", "scroll down". A 2-minute run costs about 12¢.

Changelog

v1 · Oct 7, 2026 First public release: /v1/decide, /v1/listen, /v1/status, Python and JS SDKs.