Skip to content

CAIL

Typed C++23 calls to any LLM provider. One model, one generation API.

CAIL is a C++23 SDK with a static library. Write a generation request once and run it against OpenAI, Anthropic, Gemini, OpenRouter, Mistral, Azure Foundry, Charm Hyper, Ollama Cloud, OpenCode, or your own local server. Streaming, structured outputs, tools, agents with memory, embeddings, and local file or PDF attachments use the same provider-neutral API.

your-app.cpp
auto response = cail::generate_text({
    .model = cail::openai("gpt-6-luna"),
    .prompt = "Summarize the main idea of this paragraph.",
});
std::cout << response->text;

One API

Provider-neutral by design

You write against a single generation API. The provider is a value you pass in, so switching models or vendors is a one-line change.

01

Typed requests and results

Designated initializers for requests, Result values for responses. No JSON hand-rolling and no vendor structs in your code.

02

Structured outputs

Describe a result with cail::Field<T> members. CAIL generates the JSON Schema and decodes the reply into your struct.

03

Function tools and agents

cail::tool<In, Out>() gives you typed argument decoding and a typed handler. Bundle tools with standing instructions in a cail::Agent and the tool loop, dispatch, and history are handled for you.

04

Streaming everywhere

Text, reasoning, tool-call deltas, and usage arrive as typed StreamEvents, with cancellation through std::stop_token.

05

Files, memory, and embeddings

Load images and PDFs from disk, give agents durable conversation memory, embed text for search, and use EmbeddingStore for small in-memory collections.

06

Async and coroutines

co_await generation, streaming, and embeddings, or use completion callbacks when you integrate with an existing event loop.

Providers

Hosted or on your own machine

Every provider reads its API key from the environment by default, and each adapter reports what it can encode, decode, and stream.

01

Hosted providers

OpenAI, Anthropic, Gemini, OpenRouter, Mistral, Azure Foundry, Charm Hyper, Ollama Cloud, and OpenCode services.

02

Local servers

Point cail::local at llama.cpp, Ollama, LM Studio, or vLLM. The same generation, streaming, tools, and structured output APIs.

03

OpenAI-compatible

Any Chat Completions endpoint works through cail::create_chat_completions, with an optional session header.

04

Build once, reuse

Install CAIL with CMake and link cail::cail in your application. Include the provider headers you need.