Quickstart
Get a key at megachadcua.com/app. New accounts get $2 free, no card. Then one call:
# export MEGACHAD_API_KEY=mcc_live_… curl https://api.megachadcua.com/v1/decide \ -H "Authorization: Bearer $MEGACHAD_API_KEY" \ -H "content-type: application/json" \ -d '{"schema": {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Click Pricing", "Scroll down", "Nothing yet"]}}, "text": "take me to sign in"}' # {"session_id": "dcd_…", # "fields": {"action": {"value": "Click Sign in", "confidence": 0.98, "probabilities": {…}}}, # "transcript": "take me to sign in", "latency_ms": 112, "model_ms": 15.2, "usage": {"credits": 1}}
# pip install megachad · export MEGACHAD_API_KEY=mcc_live_… import megachad d = megachad.decide({"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}}, "take me to sign in") print(d["action"].value, d["action"].confidence) # Click Sign in 0.98
// npm i megachad · Node: export MEGACHAD_API_KEY=mcc_live_… import { decide } from "megachad"; const d = await decide({ schema: { action: { question: "What should the agent do?", options: ["Click Sign in", "Scroll down", "Nothing yet"] } }, text: "take me to sign in" }); console.log(d.fields.action.value, d.fields.action.confidence); // Click Sign in 0.98
Voice works the same way: send the same schema in a start message over a WebSocket, stream the mic, read fields events. See Listen, or skip to megachad.listen_mic.
Base URL: https://api.megachadcua.com. Every request and response body is JSON. CORS is open on /v1/*.
Authentication
Keys look like mcc_live_…. Get, roll and revoke them at /app (up to 10 live keys). Send the key one of three ways:
| Where | How | Use it for |
|---|---|---|
| Header | Authorization: Bearer mcc_live_… | Servers. REST and WebSocket. |
| Header | x-api-key: mcc_live_… | Clients that can't set Authorization. |
| Subprotocol | new WebSocket(url, ["megachad", "mcc_live_…"]) | Browsers. They can't set headers on a WebSocket. |
?key= in the URL is not accepted. URLs end up in logs. You get invalid_key (401 / close 4001).
Rolling a key kills the old one at once, including open sockets: they get invalid_key and close 4001 within a minute.
A key in a public web page is a key anyone can spend. Use the subprotocol for internal tools and prototypes. For public sites, keep the key on your server and proxy /v1/listen, or call /v1/decide from your backend.
Decide (text)
POST /v1/decide. One decision from text. One model pass: about 15 ms on the GPU, 100 to 250 ms round trip.
Request
| Field | Type | What it does |
|---|---|---|
schema | object, required | The decisions to make. See Schema. Max 64 KB as JSON. |
text | string or list, required | What was said. A string, or a conversation: [{"speaker": "user", "text": "…"}, …]. Max 200,000 characters. |
speaker | string | "user" (default) or "other", when text is a string. |
timeout_ms | integer | How long to wait for the model. Default 15,000. Clamped to 1,000 to 60,000. |
Full example
curl https://api.megachadcua.com/v1/decide \
-H "Authorization: Bearer $MEGACHAD_API_KEY" \
-H "content-type: application/json" \
-d '{
"schema": {
"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Click Pricing", "Scroll down", "Nothing yet"]},
"confirm": {"question": "Did the user confirm?", "type": "yes_no"},
"urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]}
},
"text": "yes, sign me in, quick"
}'
{
"session_id": "dcd_8Hq2vT0bXk1mPz4R",
"fields": {
"action": {
"value": "Click Sign in",
"confidence": 0.98,
"probabilities": {"Click Sign in": 0.98, "Click Pricing": 0.01, "Scroll down": 0.006, "Nothing yet": 0.004}
},
"confirm": {"value": true, "confidence": 0.93},
"urgency": {"value": "high", "level": 2, "confidence": 0.81}
},
"transcript": "yes, sign me in, quick",
"latency_ms": 112,
"model_ms": 15.2,
"usage": {"credits": 1}
}
Response
fields: one entry per schema field. Alwaysvalueandconfidence(0 to 1). Options addprobabilitiesper option; scales addlevel. See field kinds.usage.credits: what this call cost. 1 for normal calls, more for big ones.latency_ms: our end, request in to answer out.model_ms: GPU time.session_id:dcd_…. Quote it when you email us.
If the fast path is down, the call falls back to a short voice-style session (~1.5 s). Same response shape, but confidence may be null and there are no probabilities. Rare.
Conversations
{"schema": {"agreed": {"question": "Did they agree on Tuesday?", "type": "yes_no"}},
"text": [{"speaker": "user", "text": "can we do Tuesday?"},
{"speaker": "other", "text": "yes, Tuesday works"}]}
Errors come back as {"error": {"code": "…", "message": "…"}} with an HTTP status. Failed calls are free. Full list in Errors.
Listen (voice)
wss://api.megachadcua.com/v1/listen. Stream audio (or text) in, get fields events the moment a decision changes. Usually mid-sentence.
The sequence
- YouOpen the socket with your key (header or subprotocol).
- YouSend
startright away, within 5 s. It carries the schema and audio format. - Server
queued, only if every model slot is busy. Repeats as you move up. Waiting in line. - Server
readywith yoursession_id. Billing starts here. Must come within 20 s or the socket closes, unbilled. - YouStream binary PCM16 audio in real time. Send
textandspeakermessages whenever. - Server
fieldswhenever a decision changes,transcriptas words land. - YouSend
endwhen there's no more input. - Server
done: final value for every field plus the full transcript. - Server
usage: seconds and credits for the session. - ServerCloses with 1000. Anything else is an error close, preceded by an
errorevent.
If the key, plan or limits say no, the socket still opens, then you get one error event and the close code (4001, 4002, 4029 or 1013). Nothing is billed.
Client messages
startfirst message, exactly once-
{ "type": "start", "schema": { "action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Go back", "Nothing yet"]}, "confirm": {"question": "Did the user confirm?", "type": "yes_no"}, "urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]} }, "audio": {"encoding": "pcm_s16le", "sample_rate": 16000}, "transcript": true, "min_change": 0.05, "speaker": "user" }audio.sample_rate: 8,000 to 48,000; 16,000 is best.transcript: falseturns off transcript events.min_change: only report a field when its confidence moves this much.speaker: who talks first. A secondstartis ignored with abad_messageerror: the schema is fixed for the session. - Audio binary frames
Raw 16-bit little-endian mono PCM at the sample rate you declared. Any frame size up to 256 KB. Send it as it's recorded: the server allows 1.5x real time after a 5 s burst, then closes with
too_fast. Silence counts as frames, so a live mic never idles out.text{"type": "text", "text": "click sign in"}Words as if heard. For typed input, or your own speech recognition.
speaker{"type": "speaker", "speaker": "other"}Who is talking from now on:
"user"or"other". See Speakers.end{"type": "end"}No more input. The server answers
done, thenusage, then closes. Closing the socket yourself works too; you just don't getdone.
Server messages
queued{"type": "queued", "position": 2, "eta_seconds": 30, "message": "Every slot is busy. You're #2 in line."}ready{"type": "ready", "session_id": "lsn_Q7c1mZp0Lr8sK2aT"}The model is listening. Billing starts.
fields{"type": "fields", "fields": {"action": {"value": "Scroll down", "confidence": 0.97}}, "latency_ms": 15.1}Only the fields that changed. Act on it when
confidenceis high enough for you; 0.85 to 0.9 is a good start.transcript{"type": "transcript", "final": "scroll down", "partial": "a bit more"}finalis settled text;partialis words still being recognized.done{"type": "done", "fields": {"action": {"value": "Scroll down", "confidence": 0.98}, "confirm": {"value": true, "confidence": 0.95}, "urgency": {"value": "high", "confidence": 0.83}}, "transcript": "scroll down, yes, it's urgent"}After
end: every field, plus the full transcript.usage{"type": "usage", "session_id": "lsn_Q7c1mZp0Lr8sK2aT", "seconds": 12, "credits": 120}Right before the socket closes, however it closes.
error{"type": "error", "code": "too_fast", "message": "Audio is arriving faster than real time for 16000 Hz. …"}Fatal ones are followed by a close.
bad_messageisn't fatal: that one message was dropped. Codes in Errors.
A plain GET /v1/listen (no upgrade) returns this protocol as JSON.
Schema & field kinds
A schema is an object. Keys are your field names; each value is one of three kinds. Same schema for /v1/decide and /v1/listen.
{
"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]},
"confirm": {"question": "Did the user confirm?", "type": "yes_no"},
"urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]}
}| Kind | Declare | Value | Extras |
|---|---|---|---|
| Options | "options": [...] | The option text, exactly as you sent it. | probabilities: option text to probability (/v1/decide). |
| Yes / no | "type": "yes_no" | true or false. On sockets, null until the conversation touches it, unless you add "unknown": false. On /v1/decide, always true or false. | |
| Scale | "scale": [...], low to high | The level text. | level: its 0-based index. |
questionis what the model answers. Optional (defaults to the field name). Write one anyway.- Options and scales can't be empty. A field with none of the three is rejected (400
bad_schemaon/v1/decide). - Give option lists an out like
"Nothing yet", so the model isn't forced to pick an action while the user is still talking. - Max 64 KB of schema JSON.
- The schema is fixed for a socket. New screen, new options: open a new session.
Speakers
Two speakers: "user" (the default, the person driving the agent) and "other" (anyone else on the call). Anything else is rejected: bad_message on sockets (the message is dropped), 400 bad_request on /v1/decide.
- Sockets: set
"speaker"instart, or send{"type": "speaker", "speaker": "other"}any time. It applies to everything after it. /v1/decide:"speaker"next to a stringtext, or a list of{"speaker", "text"}lines.- Transcripts: once
otherhas been set, transcript lines come back prefixeduser:andother:. Single-speaker sessions get no labels.
Waiting in line
Voice sessions stream, so the model has a fixed number of voice slots (see /v1/status). Text decisions take ~100 ms each and have their own, bigger pool, so open voice sessions never block /v1/decide. When a pool is full:
- Voice waits. Your socket stays open and you get
{"type": "queued", "position": 2, "eta_seconds": 30}every time your position changes. First come, first served. When a slot frees up, thestartyou already sent is replayed andreadyfollows. Audio and text sent while waiting are dropped, and waiting is never billed. After 2 minutes in line:at_capacity, close 1013. - Text doesn't queue. When the text pool is full,
/v1/decidereturns 503at_capacitywithRetry-After: 2, after at most ~2 s of retrying on our side. Not billed. Retry it. - The 5 s
startand 20 sreadyclocks start when you get a slot, not when you connect. - Free accounts never take the last free slot. Add a card and you can.
Limits
| What | Limit | Over it |
|---|---|---|
start | Within 5 s of getting a slot. Nothing may come before it. | no_start, 1008 |
ready | Within 20 s of getting a slot | no_ready, 1013, not billed |
| Idle socket | 60 s with no frames at all | idle, 1000 |
| Session length | 2 hours from ready | max_duration, 1000 |
| Audio rate | 1.5x real time of the declared sample_rate, after a 5 s burst | too_fast, 1008 |
| Audio frame | 256 KB | bad_message, frame dropped |
| Sent while the model connects | 2 MB | too_fast, 1008 |
| Text frame | 64 KB | bad_message, frame dropped |
| Text frame rate | Burst of 200, then 50 a second | too_fast, 1008 |
| Schema | 64 KB of JSON | bad_schema (socket), 413 too_large (decide) |
| Decide text | 200,000 characters | 413 too_large |
| Decide timeout | timeout_ms 1,000 to 60,000 (default 15,000) | Clamped |
| Open sessions | Voice: free credits 1, with a card 2. Text decisions in flight: free 2, with a card 8. Counted separately. | concurrency_limit, 429 / 4029 |
| New voice sessions | Free credits: 30 per 10 minutes | slow_down, 429 / 4029 |
| Waiting in line | 2 minutes | at_capacity, 1013 |
| API keys | 10 live keys per account | Revoke one at /app |
Need more than 2 at a time? Email hello@megachadcua.com.
Errors & close codes
REST errors: {"error": {"code": "…", "message": "…"}} with the HTTP status. Socket errors: an {"type": "error", "code", "message"} event, then the close code. Switch on code; messages can change.
Billed? A failed request is never billed. On a socket, time after ready is billed up to the close, whatever closes it. "Time served" below means exactly that.
| Code | HTTP / close | Meaning | Billed? | What to do |
|---|---|---|---|---|
invalid_key | 401 / 4001 | Key missing, wrong, revoked, or sent as ?key=. Mid-session: the key was rolled. | No / time served | Use a live key from /app, in a header or the subprotocol. |
payment_required | 402 / 4002 | Free credits used up and no card. Open sockets get it at the next minute's check. | No / time served | Add a card at /app. |
payment_failed | 402 / 4002 | Your last payment failed. | No | Update the card at /app. |
spend_limit | 402 / 4002 | You hit the monthly limit you set. Open sockets get it at the next minute's check. | No / time served | Raise or remove it at /app, or wait for the 1st (UTC). |
concurrency_limit | 429 / 4029 | Too many sessions open on your account. | No | Close one, or wait for it to finish. |
slow_down | 429 / 4029 | Free account opened 30+ voice sessions in 10 minutes. | No | Wait, or add a card. |
at_capacity | 503 / 1013 | Every slot is busy. Text: within ~2 s. Voice: after 2 min in line. | No | Retry after Retry-After (2 s), or reconnect. |
model_unavailable | 503 / 1013 | Model backend unreachable. | No | Retry in a few seconds. Check status. |
model_offline | 503 / 1013 | Model backend is switched off. | No | Retry later. Check status. |
no_ready | 1013 | The model didn't start within 20 s. | No | Reconnect. |
bad_request | 400 | Body isn't JSON, no schema, no text, or an unknown speaker. Also when the model rejects the input. | No | Fix the body. |
bad_schema | 400 / 1008 | Decide: a field without options, scale or "type": "yes_no", or an empty list. Socket: a schema over 64 KB. | No | Fix the schema. |
too_large | 413 | Schema over 64 KB or text over 200,000 characters. | No | Split it. |
no_start | 1008 | No start within 5 s, or audio or text before it. | No | Send start first, right after connecting. |
too_fast | 1008 | Audio over 1.5x real time, more than 50 text frames a second, or over 2 MB sent before ready. | Time served | Stream as recorded. Set audio.sample_rate right. Batch text. |
bad_message | socket stays open | One message dropped: not JSON, text frame over 64 KB, audio frame over 256 KB, unknown speaker, or a second start. | n/a | Fix that message. The session goes on. |
idle | 1000 | No frames at all for 60 s. | Time served | Keep the mic streaming (silence counts), or reconnect when there's input. |
max_duration | 1000 | 2 hours since ready. | Time served | Open a new session. |
session_lost | 1011 | The session expired on our side. | Time served | Open a new session. |
| model error | 502 / 1011 | The model failed or didn't answer in time. Decide codes: timeout, model_closed, model_error. | No / time served | Retry. |
not_found | 404 | No such endpoint. | No | Endpoints: POST /v1/decide, WS /v1/listen, GET /v1/status. |
Close codes at a glance
| Close | Means | Retry? |
|---|---|---|
| 1000 | Normal: after done, or idle / max_duration | Open a new one when you need it |
| 1008 | Your input broke a rule: bad_schema, no_start, too_fast | Not until you fix it |
| 1011 | Model or gateway failed after ready. Time served is billed. | Yes |
| 1013 | Busy or unavailable: at_capacity, model_unavailable, no_ready. Not billed. | Yes, in a few seconds |
| 4001 | invalid_key | With a valid key |
| 4002 | payment_required / payment_failed / spend_limit | After adding a card or raising the limit |
| 4029 | concurrency_limit / slow_down | After closing a session, or waiting |
Pricing math
Everything is credits. 1 credit = $0.0001. One meter, one line on a monthly Stripe invoice.
- Voice: 10 credits a second = 6¢ a minute. Wall-clock time from
readyto close, talking or not (the session holds a GPU slot). Per second, rounded up, 1 s minimum. - Text: 1 credit per decision = 10¢ per 1,000. A call over 4,000 input tokens counts as
ceil(tokens / 4000)decisions. Input tokens =ceil(UTF-8 bytes / 4)of the text plus the same for the schema JSON. - Free: failed calls, sessions that never reach
ready, time waiting in line. - New accounts: $2 free = 20,000 credits = about 33 minutes of voice, or 20,000 decisions. Spent first. After that, add a card or calls get
payment_required. - Spend limit: optional monthly cap in the console (Billing). We email at 80% and at 100%; at 100% calls get
spend_limituntil the 1st (UTC). No limit unless you set one.
| Example | Math | Credits | Cost |
|---|---|---|---|
| 90-second voice session | 90 s × 10 | 900 | 9¢ |
| 90.2-second voice session | rounds up to 91 s × 10 | 910 | 9.1¢ |
| 1,000 decide calls, normal size | 1,000 × 1 | 1,000 | 10¢ |
| One 10,000-token decide call | ceil(10,000 / 4,000) | 3 | 0.03¢ |
Socket that never got ready | not billed | 0 | 0¢ |
2 minutes in line, then at_capacity | not billed | 0 | 0¢ |
| An hour-long voice session | 3,600 s × 10 | 36,000 | $3.60 |
Every /v1/decide response carries usage.credits. Every socket ends with a usage event. Your console shows the running month.
Status endpoint
GET https://api.megachadcua.com/v1/status. No key. Cached for 10 s.
{
"model": "up",
"slots": 3,
"in_use": 1,
"free": 2,
"demo": true,
"today": {"decisions": 1834, "fastest_ms": 14.2}
}model:"up","down"or"unconfigured".slots,in_use,free: voice slots in total, taken, open.free: 0means new voice sessions wait in line. Text decisions don't use these.today: calls served since midnight UTC and the fastest model time.
History
GET https://api.megachadcua.com/v1/status/history?days=90. No key. Cached for 60 s. days is 1 to 90, default 90. This is what /status draws.
{
"days": [{"date": "2026-10-07", "up_pct": 99.65, "checks": 288}],
"last24h": [{"at": 1791374700000, "model": "up", "decide_p50_ms": 96}],
"incidents": [{"start": 1791331200000, "end": 1791332700000, "minutes": 25}]
}- We check every 5 minutes: the same probe as
/v1/status, plus the median wall time of/v1/decidecalls in those 5 minutes. days: oldest first, UTC days, today last.up_pctis the share of checks that found the model up,nullon a day with no checks.last24h: every check in the last 24 hours.decide_p50_msisnullwhen no decisions finished in that window.incidents: the latest 10 runs of failed checks in the window, newest first. Times are ms since epoch.endis the first check after that found the model up,nullwhile it's still down.
SDKs & example
Both SDKs pace audio to real time, map every error code to a typed exception with a retry hint, and read MEGACHAD_API_KEY from the environment.
Python
import megachad # Python 3.9+. Linux mic: apt install libportaudio2 schema = {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Go back"]}} # text d = megachad.decide(schema, "take me to sign in") print(d["action"].value, d.credits) # Click Sign in 1 # voice: talk, get decisions as you go for f in megachad.listen_mic(schema): print(f) # action='Go back' (0.97)
pip install megachad without [mic] for text and raw audio only. async with megachad.listen(schema) as s gives you the socket: send_audio, say, speaker, end. CLI: megachad decide --text "take me to pricing" --options "Pricing,Docs,Nothing yet".
JavaScript
// Browsers, Deno, Bun, Node 22+. Node 18/20: npm i megachad ws import { decide, listen } from "megachad"; const schema = { action: { question: "What should the agent do?", options: ["Click Sign in", "Scroll down", "Go back"] } }; const d = await decide({ schema, text: "take me to sign in" }); d.fields.action.value; // "Click Sign in" const s = listen({ schema, onFields: (e) => console.log(e.fields) }); await s.ready; // throws InvalidKey, PaymentRequired, AtCapacity… s.sendAudio(pcm16); // 16 kHz mono PCM16, paced for you const done = await s.end(); // final fields + transcript
Browser mic in one call: await listenMic({ key, schema, onFields }). It authenticates with the subprotocol, never the URL. Mind the key-in-a-web-page warning.
Example: voice-browser
Drive a real browser with your voice in ~60 lines of Python. It reads the links and buttons on screen, sends them as options, and clicks the one you mean once confidence hits 0.85.
pip install -r requirements.txt && playwright install chromium
export MEGACHAD_API_KEY=mcc_live_…
python voice_browser.py https://news.ycombinator.com
# no mic: typed commands over /v1/decide
printf 'open the top story\ngo back\nscroll down\n' | python voice_browser.py https://news.ycombinator.com --textSay "open the top story", "go back", "scroll down". A 2-minute run costs about 12¢.
Changelog
v1 · Oct 7, 2026 First public release: /v1/decide, /v1/listen, /v1/status, Python and JS SDKs.