Production AI inference, one endpoint away.
Kymo gives you a fast, OpenAI-compatible chat completions API plus a full playground for prompting, streaming, vision and structured output.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.KYMO_API_KEY,
baseURL: "https://kymo.volarishq.uk/api/public/v1",
});
const res = await client.chat.completions.create({
model: "kymo-large",
messages: [{ role: "user", content: "Explain quantum tunneling briefly." }],
stream: true,
});
for await (const chunk of res) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}Everything you need to ship
OpenAI-compatible
Point any OpenAI-compatible SDK at the Kymo base URL. Same request and response shapes, streaming included.
Curated model catalog
Kymo Large, Pro, Vision and Agent — general reasoning, multimodal and tool-using models behind one endpoint.
Vision & multimodal
Send images alongside text for OCR, captioning and visual Q&A directly from the playground or the API.
Structured JSON mode
Force strictly valid JSON responses for extraction pipelines and tool-driven workflows.
Keys with guardrails
Issue and revoke hashed API keys per project, with per-key rate limiting baked in.
Usage analytics
Requests, tokens, latency and error rates broken down by day, model and key.
Models
Call any of these with the same request body.
Kymo Large
kymo-largeGeneral-purpose workhorse. Fast, strong at reasoning and long-form text.
Kymo Pro
kymo-proOur largest open-weight model for demanding reasoning and code tasks.
Kymo Vision
kymo-visionMultimodal: image understanding, OCR, captioning and visual Q&A.
Kymo Agent
kymo-agentAgentic system with built-in web search, page visits and code execution.
Kymo Agent Mini
kymo-agent-miniSingle tool call per request, roughly 3x lower latency than Kymo Agent.
Get your first key in a minute.
Create an account, generate a key, and start streaming.