Agent
You can bundle a model, standing instructions, and a set of tools into a
cail::Agent, then run prompts against it. Each call runs the tool loop for
you: the agent sends its instructions as the first system message, offers its
tools, executes the tool calls, and returns the final response.
#include <cail/agent.hpp>#include <cail/openai.hpp>
#include <iostream>
int main() { cail::Agent agent({ .model = cail::openai("gpt-6-luna"), .instructions = "You are a concise research assistant.", }); auto response = agent.generate("Name one benefit of native C++ AI applications."); if (!response) { std::cerr << response.error().message << '\n'; return 1; } std::cout << response->text << '\n';}Build and run the example with OPENAI_API_KEY set:
cmake --build build --target cail_openai_agent./build/cail_openai_agentConfigure the agent
Section titled “Configure the agent”AgentConfig takes five fields:
model: any CAIL model, such ascail::openai("gpt-6-luna")or a model fromcreate_local. See Providers.instructions: standing instructions prepended to every request as the first system message. Leave it empty to send none.tools: the tools offered on every run. See Tools for creating typed handlers withcail::tool.memory: an optional conversation history backend. See Memory.conversation_id: the default conversation id used withmemory.
Add tools
Section titled “Add tools”struct WeatherQuery { std::string location;};
struct WeatherReport { std::string location; int temperature_fahrenheit{};};
auto weather_tool = cail::tool<WeatherQuery, WeatherReport>( "weather", "Get the weather in a location.", [](const WeatherQuery& query) { return WeatherReport{.location = query.location, .temperature_fahrenheit = 72}; });
cail::Agent agent({ .model = cail::openai("gpt-6-luna"), .instructions = "Answer questions about the weather.", .tools = {weather_tool},});
auto response = agent.generate("Fetch the weather for Paris.");When the model requests a tool call, CAIL decodes the arguments into your input
type, runs your handler, and sends the result back for the next model step. The
returned GenerationResponse contains the final text and each tool result.
Pass a full request
Section titled “Pass a full request”For history, images, or a session ID, pass a GenerationRequest instead of a
prompt string. The agent still prepends its instructions and offers its tools:
cail::GenerationRequest request{ .messages = {cail::Message{ .role = cail::MessageRole::user, .content = {cail::TextPart{.text = "What did we discuss earlier?"}}, }}, .session_id = "conversation-42",};auto response = agent.generate(request);Passing a request also bypasses agent memory: nothing is loaded and nothing is stored. See Memory to keep history for prompt calls.
Tune the loop
Section titled “Tune the loop”Both generate overloads accept ToolLoopOptions:
std::stop_source stop;auto response = agent.generate( "Fetch the weather for Paris and Tokyo.", { .max_rounds = 4, .stop_when = [](const cail::GenerationResponse& response, const cail::MiddlewareContext&) { return response.status == cail::GenerationStatus::refused; }, .stop = stop.get_token(), });max_roundscaps the follow-up model calls. The default is 8.stop_whenruns after a model step and before its requested tools execute; returningtruereturns that response without starting another step.stopcancels the active stream when you callstop.request_stop().
Stream a run
Section titled “Stream a run”agent.stream() takes the same event callback as a model stream and runs the
tool loop while events arrive:
auto response = agent.stream( "Fetch the weather for Paris.", [](const cail::StreamEvent& event) { if (const auto* delta = std::get_if<cail::TextDelta>(&event)) { std::cout << delta->text << std::flush; } }, {.max_rounds = 4});The callback receives text, reasoning, usage, and tool-call events from every step. See Streaming for the full event list.
Manage the loop yourself
Section titled “Manage the loop yourself”Agent is a convenience over the lower-level APIs. When you want to inspect or
drive each round yourself, call the model with LanguageModel::generate(), or
run cail::run_tool_loop and cail::stream_tool_loop directly as shown in
Tools.
Ask about an image or document
Section titled “Ask about an image or document”Pass a user Message when your prompt includes images or PDFs:
auto image = cail::load_image("chart.png");if (!image) { std::cerr << image.error().message << '\n'; return 1;}auto response = agent.generate(cail::Message{ .content = {cail::TextPart{.text = "Explain this chart."}, std::move(*image)},});Include <cail/loaders.hpp> and <utility>. You can include several text,
image, and PDF parts in the same message. When you configure memory and a
conversation id, successful turns save the full message and replay its
attachments in later calls. The same input works with stream and
generate_async. See Memory for a complete conversation example.
Single-message inputs accept only the user role and cannot contain tool calls or
a tool call ID. Use GenerationRequest when you manage the full history yourself.
Generate without blocking
Section titled “Generate without blocking”Call generate_async with a prompt and completion callback:
auto started = agent.generate_async( "What should I pack for Paris?", [](cail::Result<cail::GenerationResponse> result) { if (result) std::cout << result->text << '\n'; else std::cerr << result.error().message << '\n'; });if (!started) { std::cerr << started.error().message << '\n';}The callback receives the final result after any tool calls. Keep your application
running until it finishes. Text prompts and single user messages load configured
memory before starting and append successful turns before invoking the callback.
With generate_async, those memory loads and appends run asynchronously. Blocking
generate and stream use the synchronous memory methods unless your backend
implements the async overrides. Calls with an explicit GenerationRequest bypass
memory.
Pass ToolLoopOptions as the third argument to set max_rounds, stop_when, or
stop. The request owns the model, tools, and memory backend until completion;
objects captured by reference in handlers must remain valid. Avoid overlapping
prompt calls for the same conversation because they can load the same history.
See Advanced usage for callback scheduling and cancellation.
Stream or await without blocking
Section titled “Stream or await without blocking”stream_async runs the same tool loop as stream, with events delivered through
your callback and the final response in a completion callback. For coroutine-based
code, co_await agent.stream_async(prompt, on_event) returns the final
GenerationResponse after events are delivered. See Async and coroutines
for co_await agent.generate_async, task ownership, and cancellation.