# megachad cua > Everything about the megachad decision API in one file: concepts, keys, every endpoint, the schema format, errors, retries, limits, voice and batch protocols, SDKs, CLI, MCP, pricing, status and agent setup. Index: https://megachadcua.com/llms.txt. megachad cua is a decision model API. Send text or live speech plus a schema of questions; get back one calibrated answer per field, each with a confidence. The model decides in ~17 ms. Each question is one of N options, yes/no, or a scale. It runs on a fine-tuned Qwen 3.8 27B. Use it as the decision step of an agent or app: which on-screen control the user means, which queue a call goes to, whether they confirmed, how urgent a ticket is. It picks; it doesn't write text or reason out loud. Medical use: it routes and structures for clinicians. It doesn't diagnose. ## Endpoints Base URL `https://api.megachadcua.com`. JSON in, JSON out. | Endpoint | Key | What | | --- | --- | --- | | `POST /v1/decide` | yes | Make decisions from text. | | `POST /v1/decide/batch` | yes | Make up to 50 decisions in one request. | | `WS /v1/listen` | yes | Live voice (or text) in, decisions while they talk. A plain GET returns the protocol as JSON. | | `GET /v1/status` | no | Model, voice and capacity status. | | `GET /v1/status/history` | no | Uptime and latency history (90 days). | | `GET /v1/openapi.json` | no | This API as OpenAPI 3.1 (also at https://megachadcua.com/openapi.json). | ## Get a key 1. A human signs up at https://megachadcua.com/login?signup=1. $2 free credit, no card. The key (`mcc_live_...`) is on the next screen and in the console at https://megachadcua.com/app. 2. The key goes in one env var, `MEGACHAD_API_KEY`. The SDKs, the CLI and the MCP server read it. `MEGACHAD_BASE_URL` overrides the base URL (default `https://api.megachadcua.com`); only set it for local testing. Rules for coding agents: - Don't sign up for the user. Ask them to create a key and put it in `MEGACHAD_API_KEY` themselves: their shell profile, their secret manager, or a `.env` file that git ignores. - Read the key from the environment at runtime. Never hardcode it, never commit it, never print or log it, never put it in a URL or in code that ships to a browser. - Before anything writes or loads a `.env`, check that `.env` is in `.gitignore`. - If `MEGACHAD_API_KEY` is unset, stop and ask the user. Without it every call gets 401 `invalid_key`. - Keys are secret. Browser apps call megachad from their own backend. - 403 `email_unverified`: the account must confirm its email first. Ask the user to click the link in their inbox (or resend it at https://megachadcua.com/app). - Leaked key: the user rolls it at https://megachadcua.com/app. The old key stops working at once. ## Authentication Send the key one of three ways: | Where | How | Use it for | | --- | --- | --- | | Header | `Authorization: Bearer mcc_live_...` | Servers. REST and WebSocket. | | Header | `x-api-key: mcc_live_...` | Clients that can't set `Authorization`. | | WebSocket subprotocol | `new WebSocket(url, ["megachad", key])` | Browsers and Node's built-in WebSocket, which can't set headers. | `?key=` in the URL is rejected (401 `invalid_key`): URLs end up in logs. Up to 10 live keys per account. ## POST /v1/decide Text in, one decision per schema field out. One model pass: about 15 ms on the GPU, about 100 to 250 ms round trip. Base URL `https://api.megachadcua.com`. JSON in, JSON out. CORS is open on `/v1/*`. Request body: | Field | Type | Required | Notes | | --- | --- | --- | --- | | `schema` | object: field name to field | yes | Field name to field. Max 64 KB as JSON. | | `text` | string or list of `{speaker?, text}` | yes | What was said. Max 200,000 characters. | | `speaker` | `"user"` or `"other"` | no | Default `"user"`. Who said `text` when it's a string. | | `timeout_ms` | integer | no | Default 15,000. Clamped to 1,000 to 60,000. | Optional header `Idempotency-Key`: makes retries safe (see Retries and idempotency). ```bash curl https://api.megachadcua.com/v1/decide \ -H "Authorization: Bearer $MEGACHAD_API_KEY" \ -H "content-type: application/json" \ -d '{ "schema": { "action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}, "confirm": {"question": "Did the user confirm?", "type": "yes_no"}, "urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]} }, "text": "yes, sign me in, quick" }' ``` Response (200): ```json { "session_id": "dcd_8Hq2vT0bXk1mPz4R", "fields": { "action": {"value": "Click Sign in", "confidence": 0.98, "probabilities": {"Click Sign in": 0.98, "Scroll down": 0.012, "Nothing yet": 0.008}}, "confirm": {"value": true, "confidence": 0.93}, "urgency": {"value": "high", "level": 2, "confidence": 0.81} }, "transcript": "yes, sign me in, quick", "latency_ms": 112, "model_ms": 15.2, "usage": {"credits": 1} } ``` - `fields`: one entry per schema field. Always `value` and `confidence` (0 to 1). Options fields add `probabilities` (option text to probability). Scale fields add `level` (0-based index of `value`). - `usage.credits`: what the call cost. 1 for a normal call; more for inputs over 4,000 tokens (see Pricing). - `model_ms`: the model's decision time, ~17 ms. `latency_ms`: our side, request in to answer out: `model_ms` plus the network hop inside our infrastructure, ~90 ms from the US east coast. Your own round trip comes on top. - `session_id` (`dcd_...`): the request id. Same as the `x-request-id` response header and the row in the console's request logs. Quote it in support emails. - Rare fallback: if the fast path is down, the call runs as a short voice-style session (about 1.5 s). Same shape, but `confidence` may be `null` and there are no `probabilities`. Treat `null` as below any threshold. Conversations: send `text` as a list of lines. `speaker` is `"user"` (the person the agent serves, default) or `"other"` (anyone else). Anything else is 400 `bad_request`. ```bash curl https://api.megachadcua.com/v1/decide \ -H "Authorization: Bearer $MEGACHAD_API_KEY" \ -H "content-type: application/json" \ -d '{"schema": {"agreed": {"question": "Did they agree on Tuesday?", "type": "yes_no"}}, "text": [{"speaker": "user", "text": "can we do Tuesday?"}, {"speaker": "other", "text": "yes, Tuesday works"}]}' ``` Response (200): ```json {"session_id": "dcd_VTm6rGleoLUHmAjA", "fields": {"agreed": {"value": true, "confidence": 0.91}}, "transcript": "can we do Tuesday? yes, Tuesday works", "latency_ms": 96, "model_ms": 14.2, "usage": {"credits": 1}} ``` Failed calls are never billed. Errors come back as `{"error": {"code": "...", "message": "..."}}` with an HTTP status. Switch on `code`; messages can change. ## POST /v1/decide/batch Up to 50 decisions in one request. Each item is one `/v1/decide` call: same answer, same price, same request log row. Results come back in order. Request body (up to 4 MB): | Field | Type | Required | Notes | | --- | --- | --- | --- | | `schema` | object: field name to field | no | For every item without its own. | | `items` | list of `{text, id?, schema?, speaker?, timeout_ms?}` | yes | 1 to 50 items. | | `speaker` | `"user"` or `"other"` | no | Default `"user"`. Who said `text` when it's a string. | | `timeout_ms` | integer | no | Per item, and never past the batch's 90 s deadline. Default 15,000. Clamped to 1,000 to 60,000. | Each item: `{"text", "id"?, "schema"?, "speaker"?, "timeout_ms"?}`. `id` is yours (a string up to 200 characters, or a number), echoed back. An item's own `schema` wins over the shared one; one of the two is required. ```bash curl https://api.megachadcua.com/v1/decide/batch \ -H "Authorization: Bearer $MEGACHAD_API_KEY" \ -H "content-type: application/json" \ -d '{"schema": {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}}, "items": [{"id": "a", "text": "take me to sign in"}, {"id": "b", "text": "scroll down a bit"}, {"id": "c", "text": ""}]}' ``` Response (200). Item `c` failed on its own and wasn't billed: ```json { "results": [ {"id": "a", "index": 0, "ok": true, "session_id": "dcd_Hp3o1qC7v9bbJ37o", "fields": {"action": {"value": "Click Sign in", "confidence": 0.98, "probabilities": {"Click Sign in": 0.98, "Scroll down": 0.012, "Nothing yet": 0.008}}}, "latency_ms": 104, "usage": {"credits": 1}}, {"id": "b", "index": 1, "ok": true, "session_id": "dcd_8XV1oaHA1xGUzFIA", "fields": {"action": {"value": "Scroll down", "confidence": 0.97, "probabilities": {"Click Sign in": 0.01, "Scroll down": 0.97, "Nothing yet": 0.02}}}, "latency_ms": 118, "usage": {"credits": 1}}, {"id": "c", "index": 2, "ok": false, "error": {"code": "bad_request", "message": "Body needs `text`."}} ], "usage": {"credits": 2}, "latency_ms": 131 } ``` How it runs: - Partial failures: a bad item doesn't fail the batch. It comes back `{"id", "index", "ok": false, "error": {"code", "message"}}`, not billed. The whole batch gets a 4xx only for a bad key, the account (402, 403) or a malformed body (no `items`, over 50 items, a bad shared `schema`, over 4 MB). - Billing: each answered item bills like one decide call. `usage.credits` is the total. - Out of credits or over the spend limit midway: answered items stay answered; the rest get `payment_required` or `spend_limit` and aren't started. - Speed: up to 4 items at a time (2 on free credits). An item refused with `at_capacity` or `concurrency_limit` is retried with backoff before it fails. - Deadline: the batch answers within 90 s. No item starts in the last 10 s. Items that didn't start or ran out of time get `batch_timeout`, not billed: send them again in a new batch. Give your HTTP client a timeout over 90 s (120 s is safe). - `x-request-id` is the batch's id (`bat_...`). Each item's `session_id` is its own request id. - `Idempotency-Key` on a batch stores the whole response, failed items included. A retry with the same key gets it back and nothing is billed again. To rerun failed items, send them in a new batch with a new key. ## WS /v1/listen (live voice) `wss://api.megachadcua.com/v1/listen`. Stream audio (or text) in; get a `fields` event the moment a decision changes, usually mid-sentence. Billed per second of session time (see Pricing). A plain `GET https://api.megachadcua.com/v1/listen` (no upgrade) returns this protocol as JSON. Sequence: 1. Open the socket with the key (header, or subprotocol `["megachad", key]`). 2. Send `start` within 5 s. Nothing may come before it. 3. Server may send `queued` (every voice slot busy; repeats as you move up the line, up to 2 minutes, never billed). 4. Server sends `ready` with your `session_id`, within 20 s of getting a slot. Billing starts here. 5. Stream binary PCM16 audio in real time, and/or `text` and `speaker` messages. 6. Server sends `fields` (only the fields that changed) and `transcript`. 7. Send `end` when there is no more input. 8. Server sends `done` (every field plus the full transcript, about 1 s after `end`), then `usage`, then closes with 1000. A refused socket (bad key, no credits, limits) still opens, then gets one `error` event and its close code. Nothing is billed. Client messages: ```json {"type": "start", "schema": {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}}, "audio": {"encoding": "pcm_s16le", "sample_rate": 16000}, "transcript": true, "min_change": 0.05, "speaker": "user"} ``` - `start`: first message, exactly once. `audio.sample_rate` 8,000 to 48,000 (16,000 is best). `transcript: false` turns off transcript events. `min_change`: only report a field when its confidence moves this much. `speaker`: who talks first. The schema is fixed for the session; a second `start` gets a non-fatal `bad_message`. - Audio: binary frames of raw 16-bit little-endian mono PCM at the declared rate, up to 256 KB each. Send it as it's recorded: over 1.5x real time (after a 5 s burst) closes with `too_fast`. Silence counts as frames. - `{"type": "text", "text": "click sign in"}`: words as if heard. Up to 64 KB; a burst of 200, then 50 a second. - `{"type": "speaker", "speaker": "other"}`: who talks from now on, `"user"` or `"other"`. - `{"type": "end"}`: no more input. Server messages: ```json {"type": "queued", "position": 2, "eta_seconds": 30, "message": "Every slot is busy. You're #2 in line."} {"type": "ready", "session_id": "lsn_Q7c1mZp0Lr8sK2aT"} {"type": "fields", "fields": {"action": {"value": "Scroll down", "confidence": 0.97}}, "latency_ms": 15.1} {"type": "transcript", "final": "scroll down", "partial": "a bit more"} {"type": "done", "fields": {"action": {"value": "Scroll down", "confidence": 0.98}}, "transcript": "scroll down a bit more"} {"type": "usage", "session_id": "lsn_Q7c1mZp0Lr8sK2aT", "seconds": 12, "credits": 120} {"type": "warning", "code": "voice_unavailable", "message": "Speech recognition is down. Send text frames or use /v1/decide."} {"type": "error", "code": "too_fast", "message": "Audio is arriving faster than real time for 16000 Hz. ..."} ``` - `fields`: act when `confidence` is high enough for you; 0.85 to 0.9 is a good start. Its `latency_ms` is the model's time for that decision, ~17 ms. - Push-to-talk: on release, wait about 200 ms, then act on the last `fields` pick if it is 0.9 or more and hasn't changed for 300 ms. Wait for `done` when it's less sure or the words hedge ("no", "actually", "wait"). - `warning` `voice_unavailable`: speech recognition is down (also `"voice": "down"` in `GET /v1/status`). Not fatal: send `text` frames or use `/v1/decide`. If it breaks mid-session you get an `error` with the same code; the socket stays open for text. - Errors with a close code end the session; `bad_message` and `voice_unavailable` don't. - Yes/no fields are `null` on sockets until the conversation touches them, unless the field sets `"unknown": false`. Limits: idle 60 s with no frames at all closes (`idle`, 1000). Sessions last 2 hours at most (`max_duration`, 1000). Over 2 MB sent before `ready`: `too_fast`. Raw WebSocket, no SDK. Node 22+ has `WebSocket` built in; the same messages work from any WebSocket client. ```js listen-raw.mjs const base = process.env.MEGACHAD_BASE_URL ?? "https://api.megachadcua.com"; const ws = new WebSocket(base.replace(/^http/, "ws") + "/v1/listen", ["megachad", process.env.MEGACHAD_API_KEY]); const schema = { action: { question: "What should the agent do?", options: ["Click Sign in", "Scroll down", "Nothing yet"] } }; ws.onopen = () => ws.send(JSON.stringify({ type: "start", schema })); ws.onmessage = (m) => { const e = JSON.parse(m.data); if (e.type === "ready") { ws.send(JSON.stringify({ type: "text", text: "take me to sign in" })); // or binary 16 kHz PCM16 frames ws.send(JSON.stringify({ type: "end" })); } if (e.type === "done") console.log(e.fields.action.value); if (e.type === "error") console.error(e.code, e.message); }; ``` Prints `Click Sign in`. ## Schema A schema is a JSON object. Keys are your field names; each value is one field of one of three kinds. The same schema works for `/v1/decide`, `/v1/decide/batch` and `/v1/listen`. ```json { "action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}, "confirm": {"question": "Did the user confirm?", "type": "yes_no"}, "urgency": {"question": "How urgent is it?", "scale": ["low", "medium", "high"]} } ``` | Kind | Declare | `value` | Extras | | --- | --- | --- | --- | | Options | `"options": ["A", "B", ...]` | The option text, exactly as sent. | `probabilities`: option text to probability (`/v1/decide`). | | Yes / no | `"type": "yes_no"` | `true` or `false` on `/v1/decide`. On sockets `null` until the conversation touches it, unless `"unknown": false`. | | | Scale | `"scale": ["low", "medium", "high"]`, low to high | The level text. | `level`: its 0-based index. | Every field also returns `confidence`, 0 to 1. Rules, from the code: - `question` is what the model answers. Optional (defaults to the field name). Write one. - A field needs `options`, `scale` or `"type": "yes_no"`. `options` and `scale` can't be empty. Otherwise 400 `bad_schema` (for example `x: needs options, scale, or type "yes_no"`). - The schema can be at most 64 KB as JSON (65,536 characters). There is no separate cap on the number of fields or options. - Options and scale levels are strings (other values are turned into strings). - On a socket the schema is fixed for the session. New screen, new options: open a new session. On `/v1/decide` every call can carry a different schema. How to design a good schema: see "Schema design" in https://megachadcua.com/agents.md. ## Errors REST: `{"error": {"code": "...", "message": "..."}}` with the HTTP status. Sockets: `{"type": "error", "code", "message"}`, then the close code. Switch on `code`; messages can change. A failed request is never billed. On a socket, time after `ready` is billed up to the close, whatever closes it. ```bash curl https://api.megachadcua.com/v1/decide \ -H "content-type: application/json" \ -d '{"schema": {"x": {"type": "yes_no"}}, "text": "yes"}' ``` Response (401): ```json {"error": {"code": "invalid_key", "message": "Missing or invalid API key. Get yours at megachadcua.com/app."}} ``` ```bash curl https://api.megachadcua.com/v1/decide \ -H "Authorization: Bearer $MEGACHAD_API_KEY" \ -H "content-type: application/json" \ -d '{"schema": {"x": {"question": "Is it a bug?"}}, "text": "the app crashes"}' ``` Response (400): ```json {"error": {"code": "bad_schema", "message": "x: needs options, scale, or type \"yes_no\""}} ``` | Code | HTTP / close | Meaning | Retry? | What to do | | --- | --- | --- | --- | --- | | `invalid_key` | 401 / 4001 | Key missing, wrong, revoked, rolled, or sent as `?key=`. | No | Check `MEGACHAD_API_KEY`. Ask the user for a live key from https://megachadcua.com/app. | | `payment_required` | 402 / 4002 | Free credit used up and no card. | No | Tell the user to add a card at https://megachadcua.com/app. | | `payment_failed` | 402 / 4002 | The last payment failed. | No | Tell the user to update the card at https://megachadcua.com/app. | | `spend_limit` | 402 / 4002 | The account hit the monthly spend limit its owner set. | No | Tell the user; they raise it at https://megachadcua.com/app, or it resets on the 1st (UTC). | | `email_unverified` | 403 / 4003 | New account hasn't confirmed its email. | No | Ask the user to click the confirmation link. | | `bad_request` | 400 | Body isn't JSON, no `schema`, no `text`, an unknown speaker, or the model rejected the input. | No | Fix the body. | | `bad_schema` | 400 / 1008 | A field without `options`, `scale` or `"type": "yes_no"`, or an empty list. Socket: schema over 64 KB. | No | Fix the schema. | | `too_large` | 413 | Schema over 64 KB, text over 200,000 characters, or a batch body over 4 MB. | No | Split it. | | `bad_idempotency_key` | 400 | `Idempotency-Key` empty, over 255 characters, or not printable ASCII. | No | Use a UUID. | | `idempotency_mismatch` | 422 | This `Idempotency-Key` was used with a different body in the last 24 hours. | No | New request, new key. | | `idempotency_in_progress` | 409 | A call with this `Idempotency-Key` is still running. | Yes, after `Retry-After` (1 s) | You get its answer. | | `concurrency_limit` | 429 / 4029 | Too many calls or sessions in flight on the account. | Yes, with backoff | Fewer calls at once, or wait for one to finish. | | `slow_down` | 429 / 4029 | Free account opened 30+ voice sessions in 10 minutes. | Yes, later | Wait, or the user adds a card. | | `at_capacity` | 503 / 1013 | Every slot is busy. Text: after about 2 s of retrying on our side. Voice: after 2 minutes in line. | Yes, after `Retry-After` (2 s) | Retry with backoff. | | `model_unavailable` | 503 / 1013 | Model backend unreachable. | Yes | Retry in a few seconds; check `GET /v1/status`. | | `model_offline` | 503 / 1013 | Model backend switched off. | Later | Check `GET /v1/status`. | | `timeout`, `model_closed`, `model_error` | 502 / 1011 | The model failed or didn't answer in time. | Yes | Retry. | | `server_error` | 500 / batch item | Something failed on our side. | Yes | Retry; email hello@megachadcua.com with the `x-request-id` if it repeats. | | `batch_timeout` | batch item | The batch hit its 90 s limit before or while this item ran. | Yes | Send the item again in a new batch. | | `not_found` | 404 | No such endpoint. | No | Endpoints: `POST /v1/decide`, `POST /v1/decide/batch`, `WS /v1/listen`, `GET /v1/status`. | | `no_start` | 1008 | No `start` within 5 s, or audio/text before it. | After fixing | Send `start` first. | | `too_fast` | 1008 | Audio over 1.5x real time, too many text frames, or over 2 MB before `ready`. | After fixing | Stream as recorded; set `audio.sample_rate` right. | | `no_ready` | 1013 | The model didn't get ready within 20 s. Not billed. | Yes | Reconnect. | | `bad_message` | socket stays open | One message dropped: not JSON, frame too big, unknown speaker, or a second `start`. | n/a | Fix that message. | | `voice_unavailable` | socket stays open | Speech recognition is down. | n/a | Send `text` frames or use `/v1/decide`. | | `idle` | 1000 | No frames for 60 s. | Reconnect when there's input | Keep streaming (silence counts). | | `max_duration` | 1000 | 2 hours since `ready`. | Yes | Open a new session. | | `session_lost` | 1011 | The session expired on our side. | Yes | Open a new session. | Close codes: 1000 normal; 1008 your input broke a rule; 1011 model or gateway failed after `ready` (time served is billed); 1013 busy or unavailable (not billed); 4001 `invalid_key`; 4002 payment or spend limit; 4003 `email_unverified`; 4029 `concurrency_limit` or `slow_down`. ## Retries and idempotency Retry 409 `idempotency_in_progress`, 429, 500, 502, 503, timeouts and dropped connections. Never retry 400, 401, 402, 403, 404, 413 or 422: the same request fails the same way. Back off (0.5 s, 1 s, 2 s, up to 8 s, with jitter) and wait out `Retry-After` when it's there. To retry without paying twice, send an `Idempotency-Key` header: a value you make per request (a UUID), 1 to 255 printable ASCII characters, reused on every retry of that request. For 24 hours, per API key: - Done, same body: you get the first response again (status, body and `x-request-id`) with header `idempotent-replayed: true`. No model call, not billed again. - Still running: 409 `idempotency_in_progress`, `Retry-After: 1`. - Different body: 422 `idempotency_mismatch`. Send byte-identical JSON on retries. - First try failed with 402, 429 or 5xx: nothing is kept, so the retry really runs. ```bash curl https://api.megachadcua.com/v1/decide \ -H "Authorization: Bearer $MEGACHAD_API_KEY" \ -H "content-type: application/json" \ -H "Idempotency-Key: 4f1c2b9e-retry-docs-example" \ -d '{"schema": {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}}, "text": "take me to sign in"}' ``` Response (200; the same body again, with `idempotent-replayed: true`, if you run it twice): ```json {"session_id": "dcd_QLbis4vYDWZpqxl0", "fields": {"action": {"value": "Click Sign in", "confidence": 0.98, "probabilities": {"Click Sign in": 0.98, "Scroll down": 0.012, "Nothing yet": 0.008}}}, "transcript": "take me to sign in", "latency_ms": 109, "model_ms": 15.1, "usage": {"credits": 1}} ``` Python, raw HTTP with retries (`pip install requests`): ```python decide_http.py import os, random, time, uuid import requests API = os.environ.get("MEGACHAD_BASE_URL", "https://api.megachadcua.com") + "/v1/decide" RETRY = {409, 429, 500, 502, 503, 504} def decide(schema, text, retries=3): headers = {"Authorization": f"Bearer {os.environ['MEGACHAD_API_KEY']}", "Idempotency-Key": str(uuid.uuid4())} for attempt in range(retries + 1): wait = min(8, 0.5 * 2 ** attempt) * (0.5 + random.random()) try: r = requests.post(API, json={"schema": schema, "text": text}, headers=headers, timeout=20) except requests.RequestException: if attempt == retries: raise time.sleep(wait) continue if r.status_code in RETRY and attempt < retries: time.sleep(float(r.headers.get("retry-after") or wait)) continue body = r.json() if r.status_code != 200: raise RuntimeError(f"{r.status_code} {body['error']['code']}: {body['error']['message']}") return body d = decide({"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}}, "take me to sign in") print(d["fields"]["action"]["value"], d["fields"]["action"]["confidence"]) ``` Prints `Click Sign in 0.98`. JavaScript, raw `fetch` with retries (Node 18+, Deno, Bun, any server runtime): ```js decide-http.mjs const API = (process.env.MEGACHAD_BASE_URL ?? "https://api.megachadcua.com") + "/v1/decide"; const RETRY = new Set([409, 429, 500, 502, 503, 504]); const sleep = (s) => new Promise((ok) => setTimeout(ok, s * 1000)); export async function decide(schema, text, retries = 3) { const headers = { authorization: `Bearer ${process.env.MEGACHAD_API_KEY}`, "content-type": "application/json", "idempotency-key": crypto.randomUUID() }; for (let attempt = 0; ; attempt++) { const wait = Math.min(8, 0.5 * 2 ** attempt) * (0.5 + Math.random()); let r; try { r = await fetch(API, { method: "POST", headers, body: JSON.stringify({ schema, text }), signal: AbortSignal.timeout(20_000) }); } catch (e) { if (attempt >= retries) throw e; await sleep(wait); continue; } if (RETRY.has(r.status) && attempt < retries) { await sleep(Number(r.headers.get("retry-after")) || wait); continue; } const body = await r.json(); if (!r.ok) throw new Error(`${r.status} ${body.error.code}: ${body.error.message}`); return body; } } const d = await decide({ action: { question: "What should the agent do?", options: ["Click Sign in", "Scroll down", "Nothing yet"] } }, "take me to sign in"); console.log(d.fields.action.value, d.fields.action.confidence); ``` Prints `Click Sign in 0.98`. The published SDKs (0.1.0 on PyPI, 0.1.0 on npm) don't retry on their own and send no `Idempotency-Key`; wrap them as above or call the API directly. Built-in retries with an `Idempotency-Key` (`max_retries` / `maxRetries`) are coming in 0.1.1. ## Limits | What | Limit | Over it | | --- | --- | --- | | Schema | 64 KB as JSON (65,536 characters). No separate cap on fields or options. | 413 `too_large` (decide), `bad_schema` 1008 (socket) | | Decide text | 200,000 characters (all lines together) | 413 `too_large` | | Decide timeout | `timeout_ms` 1,000 to 60,000, default 15,000 | Clamped | | Batch | 50 items, 4 MB body, answers within 90 s | 400 `bad_request`, 413 `too_large`, item `batch_timeout` | | Text decisions in flight | Free credit: 2. With a card: 8. | 429 `concurrency_limit` | | Voice sessions open | Free credit: 1. With a card: 2. Counted apart from text. | 429 / 4029 `concurrency_limit` | | New voice sessions | Free credit: 30 per 10 minutes | 429 / 4029 `slow_down` | | `start` | Within 5 s of getting a slot | `no_start`, 1008 | | `ready` | Within 20 s of getting a slot | `no_ready`, 1013, not billed | | Waiting in line (voice) | 2 minutes | `at_capacity`, 1013 | | Idle socket | 60 s with no frames | `idle`, 1000 | | Session length | 2 hours from `ready` | `max_duration`, 1000 | | Audio rate | 1.5x real time after a 5 s burst | `too_fast`, 1008 | | Audio frame | 256 KB | `bad_message`, frame dropped | | Text frame | 64 KB; a burst of 200, then 50 a second | `bad_message` / `too_fast` | | API keys | 10 live keys per account | Revoke one at https://megachadcua.com/app | | Request logs | Kept 7 days; per hour, the first 300 calls keep bodies, the first 3,000 are logged | Not logged | Text and voice have separate pools: open voice sessions never block `/v1/decide`. When the text pool is full, `/v1/decide` retries for about 2 s on our side, then returns 503 `at_capacity` with `Retry-After: 2`. Free accounts never take the last free slot. Need more concurrency: email hello@megachadcua.com. ## SDKs and CLI Raw HTTP always works and needs nothing installed (see `/v1/decide` above). The SDKs add typed results, typed errors and voice streaming. All of them read `MEGACHAD_API_KEY` and default to `https://api.megachadcua.com`. What the registries serve today: | Install | Version | Has | | --- | --- | --- | | `pip install megachad` (Python 3.9+) | 0.1.0 | `decide`, `adecide`, `listen`, `listen_mic`, `mic`, typed errors, the `megachad` CLI (`decide`, `listen`) | | `uvx megachad ...` | 0.1.0 (from PyPI) | The CLI with nothing to install | | `npm i megachad` (Node 18+, browsers) | 0.1.0 | `decide`, `listen`, `listenMic`, typed errors | | `npx -y megachad-mcp` (Node 18+) | 0.1.0 | MCP server with tools `decide` and `choose` | Coming in 0.1.1 (in the source repo, not on PyPI or npm yet): `decide_batch` / `adecide_batch` / `decideBatch` and `megachad batch`, built-in retries with an `Idempotency-Key`, and `on_warning` / `onWarning` for the `voice_unavailable` warning. Don't call these yet. Until then: batch with raw HTTP (`POST /v1/decide/batch`), retries as in "Retries and idempotency". With 0.1.0, a mid-session `voice_unavailable` error ends an SDK voice session; reconnect, or send text with `/v1/decide`. ### Python ```bash pip install megachad ``` ```python decide.py import megachad d = megachad.decide( {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}}, "take me to sign in", ) print(d["action"].value, d["action"].confidence) ``` Prints `Click Sign in 0.98`. `d.fields`, `d.session_id`, `d.latency_ms`, `d.credits` and `d.raw` (the JSON) are there too. Conversations: pass a list of `{"speaker": "user" | "other", "text": "..."}` dicts or `(speaker, text)` tuples. Async: `await megachad.adecide(...)`. Keyword args: `key`, `speaker`, `timeout` (seconds, default 15), `base`. Errors raise `megachad.MegachadError` subclasses (`InvalidKey`, `PaymentRequired`, `ConcurrencyLimit`, `RateLimited`, `AtCapacity`, `TooFast`, `BadRequest`, `ModelError`, `SessionClosed`, `Timeout`) with `.code` (the API code), `.status`, `.retryable` and `.retry_after` (seconds): ```python try: d = megachad.decide(schema, text) except megachad.MegachadError as e: if e.retryable: ... # wait e.retry_after seconds, then retry else: raise ``` Voice or text over a socket: ```python listen.py import asyncio import megachad schema = {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}} async def main(): async with megachad.listen(schema) as s: await s.say("take me to sign in") # or: await s.send_audio(pcm), 16 kHz mono PCM16 await s.end() done = await s.wait_done() print(done["action"].value) asyncio.run(main()) ``` Prints `Click Sign in`. Session methods: `send_audio`, `say`, `speaker`, `end`, `wait_done`, `on_fields(cb)`, or `async for ev in s` for every event. From the microphone (needs `pip install "megachad[mic]"`; Linux also needs `apt install libportaudio2`): ```python import megachad schema = {"action": {"question": "What should the agent do?", "options": ["Click Sign in", "Scroll down", "Nothing yet"]}} for f in megachad.listen_mic(schema): print(f) # action='Click Sign in' (0.97), while the user talks. Ctrl-C stops. ``` ### CLI ```bash uvx megachad decide --text "take me to sign in" --options "Click Sign in,Scroll down,Nothing yet" ``` Prints `action='Click Sign in' (0.98) (112 ms)`. Flags: `--text` (or stdin), `--options "A,B,C"`, `--question`, `--field` (default `action`), `--schema ''` or a path to a `.json` file, `--speaker user|other`, `--json` (raw response), `--key`, `--base`. Talk instead: `uvx --from "megachad[mic]" megachad listen --options "Scroll down,Go back,Nothing yet"` (needs a mic). ### JavaScript / TypeScript ```bash npm i megachad ``` ```js decide.mjs import { decide } from "megachad"; const d = await decide({ schema: { action: { question: "What should the agent do?", options: ["Click Sign in", "Scroll down", "Nothing yet"] } }, text: "take me to sign in", }); console.log(d.fields.action.value, d.fields.action.confidence); ``` Prints `Click Sign in 0.98`. ESM only; the `.mjs` name makes `import` and top-level `await` work with no config. Options: `schema`, `text`, `key`, `speaker`, `timeout` (seconds, default 15), `baseUrl`. Errors are `MegachadError` subclasses with `code`, `status`, `retryable`, `retryAfter`. ```js listen.mjs import { listen } from "megachad"; const s = listen({ schema: { action: { question: "What should the agent do?", options: ["Click Sign in", "Scroll down", "Nothing yet"] } }, onFields: (e) => console.log("live:", e.fields.action?.value), }); await s.ready; s.say("take me to sign in"); // or s.sendAudio(pcm16): 16 kHz mono PCM16, paced for you const done = await s.end(); console.log(done.fields.action.value); ``` Prints `Click Sign in` last. Node 22+, browsers, Deno and Bun use the built-in WebSocket; Node 18 and 20 also need `npm i ws`. In a browser, never ship an `mcc_live_` key: keep it on your server and proxy `/v1/listen`, or call `/v1/decide` from your backend. ## MCP server `npx -y megachad-mcp` (version 0.1.0 on npm, Node 18+) is a stdio MCP server with two tools: - `choose` `{question, options, text}`: one question, one pick. Returns `Chosen: "Click Sign in" (confidence 0.98) [1 credit, 112 ms]` plus `{"choice", "confidence"}` as JSON. - `decide` `{schema, text}`: one decision per field, same schema as `/v1/decide`. Returns `action = "Click Sign in" (0.98) [1 credit, 112 ms]` plus the fields as JSON. Each call is one `/v1/decide` call. Errors (bad key, no credit, busy model, bad schema) come back as tool errors that say what to fix. Env: `MEGACHAD_API_KEY` (required), `MEGACHAD_BASE_URL` (optional). Claude Code (stores the key in your local, uncommitted `~/.claude.json`): ```bash claude mcp add megachad -e MEGACHAD_API_KEY=$MEGACHAD_API_KEY -- npx -y megachad-mcp ``` Claude Code, shared with a team through a committed `.mcp.json` (each person's own key comes from their environment): ```json { "mcpServers": { "megachad": { "type": "stdio", "command": "npx", "args": ["-y", "megachad-mcp"], "env": { "MEGACHAD_API_KEY": "${MEGACHAD_API_KEY}" } } } } ``` Cursor: `~/.cursor/mcp.json` (all projects) or `.cursor/mcp.json` (this project): ```json { "mcpServers": { "megachad": { "type": "stdio", "command": "npx", "args": [ "-y", "megachad-mcp" ], "env": { "MEGACHAD_API_KEY": "${env:MEGACHAD_API_KEY}" } } } } ``` Claude Desktop: Settings, Developer, Edit Config (`claude_desktop_config.json`, a local file); put the key in place of `mcc_live_...`: ```json { "mcpServers": { "megachad": { "command": "npx", "args": ["-y", "megachad-mcp"], "env": { "MEGACHAD_API_KEY": "mcc_live_..." } } } } ``` ## Pricing Everything is billed in credits. 1 credit = $0.0001. One metered line on a monthly Stripe invoice. No seats, no tiers. | What | Price | In credits | | --- | --- | --- | | Text decisions, `/v1/decide` and each batch item | 10¢ per 1,000 | 1 per decision | | Voice, `/v1/listen` | 6¢ per minute | 10 per second | | New accounts | $2 free, no card | 20,000 credits, about 20,000 decisions or 33 minutes of voice | - Text: a call whose input is over 4,000 tokens counts as `ceil(tokens / 4,000)` decisions. Input tokens = `ceil(UTF-8 bytes / 4)` of the text plus the same for the schema JSON. `usage.credits` on every response says what the call cost. - Voice: wall-clock time from `ready` to close, talking or not (the session holds a GPU slot). Per second, rounded up, 1 s minimum. Each session ends with a `usage` event. - Free: failed calls, sessions that never reach `ready`, time waiting in line, and voice sessions where speech recognition was down and no words came back. - After the free credit: the user adds a card at https://megachadcua.com/app, or calls get 402 `payment_required`. - Spend limit: optional monthly cap set in the console. Email at 80% and 100%; at 100% calls get 402 `spend_limit` until the 1st (UTC). Worked examples: | Example | Math | Credits | Cost | | --- | --- | --- | --- | | 1,000 normal decide calls | 1,000 x 1 | 1,000 | 10¢ | | One 10,000-token decide call | ceil(10,000 / 4,000) x 1 | 3 | 0.03¢ | | 90-second voice session | 90 x 10 | 900 | 9¢ | | 90.2-second voice session | 91 x 10 | 910 | 9.1¢ | | One hour of voice | 3,600 x 10 | 36,000 | $3.60 | | 1,000 voice minutes + 100,000 decisions | 60,000 x 10 + 100,000 x 1 | 700,000 | $70 | | Socket that never got `ready` | not billed | 0 | 0¢ | ## Status endpoints The page at https://megachadcua.com/status draws live data from two public JSON endpoints. No key needed. Read them directly; they are the status. ```bash curl https://api.megachadcua.com/v1/status ``` Response (200, cached 10 s): ```json {"model": "up", "voice": "up", "slots": 3, "in_use": 1, "free": 2, "demo": true, "today": {"decisions": 1834, "fastest_ms": 14.2}} ``` - `model`: the text model behind `/v1/decide`. `"up"`, `"down"` or `"unconfigured"`. - `voice`: speech recognition (audio on `/v1/listen`). `"up"`, `"down"` or `"unknown"` (not checked yet). Checked every 5 minutes and by live sessions. It can be `"down"` while `model` is `"up"`: then send text (`/v1/decide`, or `text` frames on the socket). Always `"down"` while `model` is. - `slots`, `in_use`, `free`: voice slots in total, taken, open. `free: 0` means new voice sessions wait in line. Text decisions have their own pool and don't use these. `slots` and `free` are `null` when there's no cap. - `demo`: whether the keyless demo on the site is on. - `today`: calls served since midnight UTC and the fastest model time in ms (`null` before the first). - Other keys may appear; ignore keys you don't know. Before a burst of work, an agent can check `model == "up"` (and `voice` for audio). Don't poll it to decide whether to retry a single call; retry with backoff instead. ```bash curl "https://api.megachadcua.com/v1/status/history?days=2" ``` Response (200, cached 60 s): ```json { "days": [{"date": "2026-10-09", "up_pct": 100, "checks": 288}, {"date": "2026-10-10", "up_pct": 99.65, "checks": 288}], "last24h": [{"at": 1791374700000, "model": "up", "voice": "up", "decide_p50_ms": 96}], "incidents": [{"start": 1791331200000, "end": 1791332700000, "minutes": 25}], "voice": {"days": [{"date": "2026-10-09", "up_pct": 100, "checks": 288}, {"date": "2026-10-10", "up_pct": 100, "checks": 288}], "incidents": []} } ``` - `days` (1 to 90, default 90): oldest first, UTC days, today last. `up_pct` is the share of 5-minute checks that found the model up; `null` on a day with no checks. - `last24h`: every check in the last 24 hours. `decide_p50_ms`: median wall time of `/v1/decide` calls in that 5-minute window, `null` when none finished. - `incidents`: the latest 10 runs of failed checks, newest first. Times are ms since the epoch. `end` is `null` while it's still down. - `voice`: the same for speech recognition. Incidents as RSS: https://megachadcua.com/status.rss. ## Agents Two kinds of agents use megachad. Coding agents (Claude Code, Codex, Cursor, Gemini CLI, Aider, Windsurf, Copilot) add it to a project. Runtime agents (a computer-use agent, a voice bot, a triage agent) call it as their decision step. Both start from https://megachadcua.com/llms.txt. ### Coding agents One-line prompt for any coding agent: ```text Read https://megachadcua.com/llms.txt and add megachad to this project: . Use the API key from the MEGACHAD_API_KEY env var; if it's unset, ask me. Never hardcode or commit the key. Call the HTTP API directly unless the project already uses a megachad SDK. ``` What the agent should do: 1. Check the key: `test -n "$MEGACHAD_API_KEY"`. If it's unset, stop and ask the user to sign up at https://megachadcua.com/login?signup=1 and set `MEGACHAD_API_KEY`. Never hardcode, commit, print or log it. 2. Prove the key works with one call (the curl under "Runtime agents" below). 401: wrong key. 403 `email_unverified`: the user confirms their email. 402: the user adds a card. 3. Write the integration. Raw HTTP is the default: no dependency, nothing to break (https://megachadcua.com/docs.md, "Retries and idempotency", has Python and JavaScript versions with retries). Use the SDK if the project already depends on it or needs live voice. 4. Design the schema with the rules under "Schema design". 5. Act on an answer only above a confidence threshold; below it, confirm with the user or hand off to a human. 6. In tests, mock the HTTP call. Real calls cost credits; run them only when the user asks. #### Claude Code Install the skill (Claude Code loads it when a task matches; it also shows up as `/megachad`): ```bash mkdir -p ~/.claude/skills/megachad && curl -fsSL https://megachadcua.com/agents/megachad/SKILL.md -o ~/.claude/skills/megachad/SKILL.md ``` For one repo only, put it in `.claude/skills/megachad/SKILL.md` instead and commit it. The skill: https://megachadcua.com/agents/megachad/SKILL.md. MCP server, so Claude can call megachad itself (`choose` and `decide` tools): ```bash claude mcp add megachad -e MEGACHAD_API_KEY=$MEGACHAD_API_KEY -- npx -y megachad-mcp ``` #### Codex and any repo: AGENTS.md Append the megachad block to the repo's `AGENTS.md`: ```bash curl -fsSL https://megachadcua.com/agents/AGENTS.md >> AGENTS.md ``` Codex, Cursor, GitHub Copilot's coding agent, Windsurf, Jules, Zed and others read `AGENTS.md` (list: https://agents.md). Gemini CLI reads it with `{"context": {"fileName": "AGENTS.md"}}` in `.gemini/settings.json`. Aider reads it with `read: AGENTS.md` in `.aider.conf.yml`. The block: https://megachadcua.com/agents/AGENTS.md. #### Cursor Project rule (type "Apply Intelligently": Cursor pulls it in when the task matches its description): ```bash mkdir -p .cursor/rules && curl -fsSL https://megachadcua.com/agents/megachad.mdc -o .cursor/rules/megachad.mdc ``` MCP server: add this to `~/.cursor/mcp.json` or `.cursor/mcp.json`. Cursor fills `${env:MEGACHAD_API_KEY}` from your environment, so the file holds no key: ```json { "mcpServers": { "megachad": { "type": "stdio", "command": "npx", "args": [ "-y", "megachad-mcp" ], "env": { "MEGACHAD_API_KEY": "${env:MEGACHAD_API_KEY}" } } } } ``` ### Runtime agents Use megachad as the decision step: the agent has a fixed set of possible next moves, a person said or wrote something, and the agent needs to know which move they mean, fast, with a number that says how sure. ```bash curl https://api.megachadcua.com/v1/decide \ -H "Authorization: Bearer $MEGACHAD_API_KEY" \ -H "content-type: application/json" \ -d '{"schema": {"action": {"question": "Which on-screen control does the user want the agent to use?", "options": ["Sign in", "Pricing", "Docs", "Nothing yet", "Not sure"]}}, "text": "take me to sign in"}' ``` Response (200): ```json {"session_id": "dcd_7KFTWwTLLOmXjo5N", "fields": {"action": {"value": "Sign in", "confidence": 0.97, "probabilities": {"Sign in": 0.97, "Pricing": 0.006, "Docs": 0.004, "Nothing yet": 0.012, "Not sure": 0.008}}}, "transcript": "take me to sign in", "latency_ms": 104, "model_ms": 15.3, "usage": {"credits": 1}} ``` A computer-use agent's step: the controls on screen become the options, plus two ways out. ```python cua_step.py import os import requests API = os.environ.get("MEGACHAD_BASE_URL", "https://api.megachadcua.com") + "/v1/decide" THRESHOLD = 0.85 # tune it on your own cases (console: Evals) def next_action(utterance, controls): schema = {"action": {"question": "Which on-screen control does the user want the agent to use?", "options": controls + ["Nothing yet", "Not sure"]}} r = requests.post(API, headers={"Authorization": f"Bearer {os.environ['MEGACHAD_API_KEY']}"}, json={"schema": schema, "text": utterance}, timeout=20) r.raise_for_status() f = r.json()["fields"]["action"] if f["value"] == "Nothing yet": return ("wait", None) # the user is still talking, or asked for nothing if f["value"] in controls and (f["confidence"] or 0) >= THRESHOLD: return ("click", f["value"]) return ("ask", f["value"]) # "Not sure" or low confidence: confirm with the user print(next_action("take me to sign in", ["Sign in", "Pricing", "Docs"])) ``` Prints `('click', 'Sign in')`. #### Schema design - One question per field. "Which team, and how urgent?" is two fields, `team` and `urgency`, answered in one call. - Options are short, distinct and mutually exclusive. Use the label the user sees ("Sign in", not `btn_auth_2`); the value comes back verbatim, so map it to your ids with a dict. - Always add a way out. "Nothing yet" for "the user hasn't asked for anything" (mid-sentence, small talk); "Not sure" or "Other" for "none of these fit". Without one the model must pick one of your actions. - Use `"type": "yes_no"` for binary questions ("Did the user confirm?") and `"scale"` for ordered levels, listed low to high. - Write the `question` as the decision the agent faces: "Which on-screen control does the user want the agent to use?", "Which queue should this call go to?". - Options can change on every `/v1/decide` call: send what's on screen now. On a socket the schema is fixed; open a new session when the screen changes. - Two speakers: send `text` as `[{"speaker": "user", "text": ...}, {"speaker": "other", "text": ...}]`. - No cap on fields or options besides the 64 KB schema limit. Input over 4,000 tokens costs more than one decision. #### Confidence and thresholds - `confidence` is calibrated, 0 to 1. Pick a threshold per field: at or above it, act; below it, confirm ("Did you mean Sign in?") or hand off. - Start at 0.85 to 0.9 for actions. Then measure: the console's Evals (https://megachadcua.com/app) runs your own cases and recommends the lowest threshold that is right at least 97% of the time on the cases it answers. - Read `probabilities` (options fields) for the runner-up: when the top two are close, ask between them. - `confidence` can be `null` (rare fallback path): treat it as below threshold. - Destructive, irreversible, financial, medical or safety-critical actions need a human to confirm, whatever the confidence. Medical: megachad routes and structures for a clinician to review; it doesn't diagnose. #### Batching Many independent texts with the same questions (a ticket queue, a dataset): `POST /v1/decide/batch`, up to 50 items per request, results in order, billed per answered item. Resend items that come back `batch_timeout` or with a retryable error (`at_capacity`, `concurrency_limit`, `server_error`, `timeout`, `model_closed`, `model_error`) in a new batch. ```python triage_batch.py import os import requests API = os.environ.get("MEGACHAD_BASE_URL", "https://api.megachadcua.com") + "/v1/decide/batch" RETRYABLE = {"batch_timeout", "at_capacity", "concurrency_limit", "server_error", "timeout", "model_closed", "model_error"} schema = {"team": {"question": "Which team should handle this ticket?", "options": ["Billing", "Bug", "Account", "Other"]}} def decide_all(texts, size=50): out, todo = [None] * len(texts), list(range(len(texts))) for _ in range(3): retry = [] for i in range(0, len(todo), size): items = [{"id": n, "text": texts[n]} for n in todo[i:i + size]] r = requests.post(API, headers={"Authorization": f"Bearer {os.environ['MEGACHAD_API_KEY']}"}, json={"schema": schema, "items": items}, timeout=120) r.raise_for_status() for res in r.json()["results"]: if res["ok"]: out[res["id"]] = res["fields"]["team"]["value"] elif res["error"]["code"] in RETRYABLE: retry.append(res["id"]) else: out[res["id"]] = "error: " + res["error"]["code"] if not retry: break todo = retry return out tickets = ["I was charged twice this month", "the app crashes when I log in", "how do I change my email?"] for ticket, team in zip(tickets, decide_all(tickets)): print(team, "|", ticket) ``` #### Voice Live speech goes over `WS /v1/listen` (https://megachadcua.com/docs.md, "WS /v1/listen"): stream 16 kHz PCM16, get `fields` events while the person talks. Act on a `fields` event when `confidence` is 0.9 or more and the pick has held for 300 ms; wait for `done` when the words hedge ("no", "actually", "wait"). #### As a tool for an LLM agent Give the model one tool that wraps `/v1/decide`. The tool input is the request body: send it unchanged and return `fields` as the tool result. Anthropic (Messages API `tools`): ```json { "name": "megachad_decide", "description": "Make fast, calibrated decisions about a piece of text with megachad (the model decides in ~17 ms). Give a schema of fields; each field is one question answered by picking one of its options, yes/no, or a level on a scale. Returns each field's value and a confidence from 0 to 1. Use it to map what a person said to one of a fixed set of actions, intents, queues or labels. Below about 0.85 confidence, ask the person instead of acting.", "input_schema": { "type": "object", "properties": { "schema": { "type": "object", "description": "Field name to field. One decision per field. Each field has a question and exactly one of: options (pick one), type yes_no, or scale (ordered levels, low to high).", "additionalProperties": { "type": "object", "properties": { "question": { "type": "string", "description": "The question the model answers about the text." }, "options": { "type": "array", "items": { "type": "string" }, "description": "Pick one of these. The answer is the option text, verbatim. Include a way out like \"Nothing yet\" or \"Not sure\"." }, "type": { "type": "string", "enum": [ "yes_no" ], "description": "yes_no: the answer is true or false." }, "scale": { "type": "array", "items": { "type": "string" }, "description": "Ordered levels, low to high. The answer is a level, plus its 0-based index." } } } }, "text": { "type": "string", "description": "What the person said or wrote: an utterance, a transcript, a ticket." } }, "required": [ "schema", "text" ] } } ``` OpenAI (Chat Completions `tools`; the Responses API takes the same `name`, `description` and `parameters` without the `function` wrapper): ```json { "type": "function", "function": { "name": "megachad_decide", "description": "Make fast, calibrated decisions about a piece of text with megachad (the model decides in ~17 ms). Give a schema of fields; each field is one question answered by picking one of its options, yes/no, or a level on a scale. Returns each field's value and a confidence from 0 to 1. Use it to map what a person said to one of a fixed set of actions, intents, queues or labels. Below about 0.85 confidence, ask the person instead of acting.", "parameters": { "type": "object", "properties": { "schema": { "type": "object", "description": "Field name to field. One decision per field. Each field has a question and exactly one of: options (pick one), type yes_no, or scale (ordered levels, low to high).", "additionalProperties": { "type": "object", "properties": { "question": { "type": "string", "description": "The question the model answers about the text." }, "options": { "type": "array", "items": { "type": "string" }, "description": "Pick one of these. The answer is the option text, verbatim. Include a way out like \"Nothing yet\" or \"Not sure\"." }, "type": { "type": "string", "enum": [ "yes_no" ], "description": "yes_no: the answer is true or false." }, "scale": { "type": "array", "items": { "type": "string" }, "description": "Ordered levels, low to high. The answer is a level, plus its 0-based index." } } } }, "text": { "type": "string", "description": "What the person said or wrote: an utterance, a transcript, a ticket." } }, "required": [ "schema", "text" ] } } } ``` Example tool input, which is also a valid `/v1/decide` body: ```json { "schema": { "team": { "question": "Which team should handle this ticket?", "options": [ "Billing", "Bug", "Account", "Other" ] }, "urgent": { "question": "Does the customer need an answer today?", "type": "yes_no" } }, "text": "I was charged twice this month and need the refund before rent is due on Friday" } ``` Handler: ```python import os import requests def run_megachad_decide(tool_input): r = requests.post(os.environ.get("MEGACHAD_BASE_URL", "https://api.megachadcua.com") + "/v1/decide", headers={"Authorization": f"Bearer {os.environ['MEGACHAD_API_KEY']}"}, json=tool_input, timeout=20) body = r.json() if r.status_code != 200: return {"error": body["error"]} # the model reads code and message; 400 bad_schema means fix the schema return body["fields"] ```