API reference
Streaming
Read tokens as they are generated.
Streaming returns the answer as it is generated, which cuts perceived latency from seconds to the time to the first token. The transport is server-sent events over the same endpoint.
Turn streaming on
Set stream: true in the request body. Everything else about the request is unchanged.
The event stream
Each frame is a data: line holding one JSON chunk. Chunks carry a delta rather than a whole message, and the stream ends with a literal [DONE].
data: {"id":"req_8fa21c0b4e9d7a35f16c2be4","object":"chat.completion.chunk","model":"anthropic/claude-sonnet-5","choices":[{"index":0,"delta":{"role":"assistant"}}]} data: {"id":"req_8fa21c0b4e9d7a35f16c2be4","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"HTTP caching"}}]} data: {"id":"req_8fa21c0b4e9d7a35f16c2be4","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":" stores responses"}}]} data: {"id":"req_8fa21c0b4e9d7a35f16c2be4","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":18,"completion_tokens":64,"total_tokens":82}} data: [DONE]The final content chunk carries finish_reason and usage, so you can record token counts without a second request.
Consuming a stream
The official SDKs expose the stream as an async iterator, so you rarely parse frames yourself.
const stream = await client.chat.completions.create({ model: "anthropic/claude-sonnet-5", messages: [{ role: "user", content: "Explain HTTP caching in two sentences." }], stream: true,}); let answer = "";for await (const chunk of stream) { answer += chunk.choices[0]?.delta?.content ?? "";}Buffer for the user, not for yourself
Render tokens as they arrive, but keep a complete copy of the text. Chunk boundaries fall inside words and sentences, so anything that parses the answer should run on the finished string.
Cancelling
Abort the request to stop generation. Odyssey stops the upstream call and bills only the tokens produced up to that point, which show up in Activity with a cancelled finish reason.
const controller = new AbortController(); const stream = await client.chat.completions.create( { model: "anthropic/claude-sonnet-5", messages, stream: true }, { signal: controller.signal },); // Stop generating, for example when the reader navigates away.controller.abort();Failures mid-stream
Failures before the first token are ordinary HTTP errors. Once tokens have been sent the status code is already 200, so a failure arrives as an error frame and the stream ends.
data: {"error":{"type":"upstream_error","code":"stream_interrupted","message":"The model stopped responding partway through.","request_id":"req_3b9e1c7a0d24f5e86c1a2b40"}} data: [DONE]Treat a stream that ends without finish_reason as incomplete. Retrying sends the whole prompt again, so decide whether a partial answer is better than a second charge.