Microsoft.Extensions.AI integration (IChatClient)

LiteRtLmSharp.Extensions.AI is a separate, optional companion package that exposes a LiteRtLmSharp on-device model as a Microsoft.Extensions.AI.IChatClient. IChatClient is the .NET ecosystem's provider-agnostic abstraction for chat models, so this one package makes the model usable from:

This is the foundation: it depends only on lightweight abstractions (Microsoft.Extensions.AI.Abstractions + Microsoft.Extensions.DependencyInjection.Abstractions for the DI helpers), and the Semantic Kernel connector is built on top of it.

Package Depends on Target
LiteRtLmSharp.Extensions.AI LiteRtLmSharp (same version) + Microsoft.Extensions.AI.Abstractions 10.7.0 + Microsoft.Extensions.DependencyInjection.Abstractions 10.0.9 net10.0
<PackageReference Include="LiteRtLmSharp" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.Extensions.AI" Version="1.1.1" />

Quick start

using LiteRtLmSharp;
using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;

using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
    ModelPath = "gemma-4-E2B-it.litertlm",
    Backend = LiteRtBackend.Cpu,
    MaxNumTokens = 4096,
});

using IChatClient client = new LiteRtChatClient(engine, modelId: "gemma-4-E2B-it");

// Blocking:
ChatResponse response = await client.GetResponseAsync("Write one upbeat sentence about on-device AI.");
Console.WriteLine(response.Text);

// Streaming:
await foreach (ChatResponseUpdate update in client.GetStreamingResponseAsync("Tell me a joke"))
    Console.Write(update.Text);

Dependency injection

// Container loads, owns and disposes a single shared engine (one engine per process):
services.AddLiteRtChatClient(new LiteRtEngineOptions
{
    ModelPath = "gemma-4-E2B-it.litertlm", Backend = LiteRtBackend.Cpu, MaxNumTokens = 4096,
});

// ... or register over an engine you already loaded (you dispose it):
services.AddLiteRtChatClient(engine, modelId: "gemma-4-E2B-it");

// Elsewhere:
IChatClient client = serviceProvider.GetRequiredService<IChatClient>();

eager: true on the options overload loads the weights at registration (so a bad model path/backend throws there) instead of lazily on first use.

Use with Microsoft Agent Framework

MAF builds agents over any IChatClient — there is no separate "MAF connector" to write:

using Microsoft.Agents.AI;

var agent = new ChatClientAgent(client, instructions: "You are a concise, helpful assistant.");
AgentRunResponse reply = await agent.RunAsync("What can you do?");

Use with the Microsoft.Extensions.AI pipeline

Because it's a standard IChatClient, it composes with the ecosystem's middleware (from the Microsoft.Extensions.AI package) — caching, telemetry, logging, and function invocation:

IChatClient pipeline = client.AsBuilder()
    .UseFunctionInvocation()      // auto-invokes AIFunctions returned in ChatOptions.Tools
    // .UseDistributedCache(...)  // caching, telemetry, logging, … all compose here
    .Build();

Function calling (tools)

Function tools in ChatOptions.Tools are passed to the model; the model's tool calls are surfaced as FunctionCallContent (with ChatFinishReason.ToolCalls). Compose UseFunctionInvocation() and the pipeline runs the loop: the model emits a call → the matching AIFunction is invoked → its result is returned to the model → the model answers from it.

using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;

AIFunction getWeather = AIFunctionFactory.Create(
    (string city) => $"22°C and sunny in {city}",
    name: "get_weather", description: "Gets the current weather for a city.");

IChatClient client = new LiteRtChatClient(engine)
    .AsBuilder()
    .UseFunctionInvocation()      // drives the tool loop
    .Build();

var options = new LiteRtChatOptions
{
    Tools = [getWeather],
    EnableConstrainedDecoding = true,   // recommended (see below); off by default
};

ChatResponse response = await client.GetResponseAsync("What's the weather in Paris?", options);
Console.WriteLine(response.Text);       // "The weather in Paris is 22°C and sunny."

Constrained decoding. Set LiteRtChatOptions.EnableConstrainedDecoding = true so the model emits valid, schema-shaped tool-call arguments — strongly recommended for small on-device models. It is off by default and works on every platform (the linux-x64 restriction of earlier releases is gone: the official LiteRT-LM v0.16.0 prebuilts embed the constraint provider). Tools still work without it; arguments are just not grammar-constrained. On a plain ChatOptions, set the enable_constrained_decoding key in AdditionalProperties instead.

Tool choice (ChatOptions.ToolMode). The native API has no tool_choice parameter (the model always decides), so the connector emulates the MEAI modes as best it can:

ToolMode Behavior
Auto (default) All tools are offered; the model decides whether to call one.
None No tools are offered, so the model can't call any.
RequireAny All tools offered + a system-prompt instruction to call one — best-effort.
RequireSpecific(name) Only that tool is offered, plus the instruction naming it — best-effort.

RequireAny / RequireSpecific are best-effort, not a guarantee: unlike a cloud API's server-enforced tool_choice: "required", the connector can only instruct the on-device model (and narrow the tool list) — the decoder isn't forced. In practice a capable model calls the tool when the request is plausibly related to it, but may ignore the instruction for a clearly-unrelated prompt. (Semantic Kernel's FunctionChoiceBehavior.Auto()/None()/Required() map to these same modes.)

The native LiteRtLmSharp tools API is still available for full control (constrained decoding, custom tool-call parsing); the chat client is the MEAI-idiomatic path on top of it.

Structured output (ResponseFormat)

A JSON-schema ChatOptions.ResponseFormat is enforced during sampling (native LiteRT-LM v0.15.0+): the client arms the LlGuidance constraint provider on the conversation and attaches the schema as a per-send constraint, so every token the model emits is masked to schema-conforming continuations — the reply is guaranteed to be a JSON document matching the schema, not just nudged toward it.

ChatResponse r = await client.GetResponseAsync(messages, new ChatOptions
{
    ResponseFormat = ChatResponseFormat.ForJsonSchema(mySchemaJsonElement),
});
// r.Text parses as JSON conforming to mySchemaJsonElement — enforced, not prompted.

Rules and caveats:

  • Schema-less ChatResponseFormat.Json ("JSON mode") is not enforced — without a schema there is nothing precise to constrain, so it keeps the previous prompt-driven behavior.
  • Not combinable with Tools on one request. The schema masks every generated token, which makes emitting a tool call impossible, so a request carrying both throws ArgumentException. Run the tool phase first, then request the schema-formatted answer in a separate call. In stateful mode both phases can share one conversation: set ConstraintProvider on the client's conversation-options template, run the tool request without a ResponseFormat, then send the schema-formatted request as a continuation of the same ConversationId — the reply is schema-enforced and informed by the tool results (covered by a model test).
  • Stateful mode: the schema must be present on the conversation's first call (the constraint provider is fixed at creation). A schema arriving only on a continuation throws ArgumentException and leaves the conversation resumable; to make every conversation schema-capable, set LiteRtConversationOptions.ConstraintProvider = LiteRtConstraintProvider.LlGuidance on the client's conversation-options template.
  • When a schema request is active, the tool-calling enable_constrained_decoding knob is dropped for that conversation (the two constrained-decoding modes are mutually exclusive natively, and the tool path is meaningless without tools).

Banning repeats and tokens: NoRepeatNgramSize / SuppressTokens

ChatOptions.FrequencyPenalty / PresencePenalty already map to the native repetition penalties. Two more logit processors have no MEAI property, so LiteRtChatOptions exposes them, backed by the no_repeat_ngram_size and suppress_tokens keys in AdditionalProperties (a plain ChatOptions with those keys works too, including options deserialized from JSON/YAML):

int[] banned = [.. engine.Tokenize(" Paris"), .. engine.Tokenize("Paris")];   // both surface forms
var options = new LiteRtChatOptions
{
    MaxOutputTokens = 128,
    NoRepeatNgramSize = 3,      // never repeat a 3-gram already produced in this reply
    SuppressTokens = banned,    // these ids can never be sampled
};
  • NoRepeatNgramSize bans any n-gram of that many tokens the reply already produced (whole-reply window). Useful when a small model echoes a prompt template verbatim. Zero or negative is rejected.
  • SuppressTokens forces the listed ids' logits to -inf on every step. Find ids with LiteRtEngine.Tokenize; most words tokenize differently with and without a leading space, so ban both. Negative ids are rejected; ids outside the vocabulary are ignored natively.
  • Both apply per request and require native LiteRT-LM v0.15.0+. The Semantic Kernel connector exposes the same two knobs on LiteRtPromptExecutionSettings.

Reasoning ("thinking")

Enable the model's reasoning mode with LiteRtChatOptions — a ChatOptions subtype that adds the knobs MEAI has no typed property for (EnableThinking, EnableConstrainedDecoding, NoRepeatNgramSize, SuppressTokens, each backed by a key in AdditionalProperties). This mirrors the Semantic Kernel connector's LiteRtPromptExecutionSettings:

var options = new LiteRtChatOptions { MaxOutputTokens = 512, EnableThinking = true };
ChatResponse r = await client.GetResponseAsync(messages, options);

(LiteRtChatOptions is a ChatOptions, so it works anywhere one is accepted — ChatOptions already exposes Temperature/TopP/TopK/MaxOutputTokens/Seed. Setting the enable_thinking key directly on ChatOptions.AdditionalProperties is equivalent.)

The reasoning trace is surfaced as a TextReasoningContent on the response (and as reasoning updates when streaming). It is excluded from ChatResponse.Text, so the answer stays clean while the reasoning stays accessible:

string answer = r.Text;
string? reasoning = r.Messages.SelectMany(m => m.Contents).OfType<TextReasoningContent>().FirstOrDefault()?.Text;

The reasoning shares the MaxOutputTokens budget with the answer. Give thinking models headroom — with too small a budget the reasoning can consume it and the answer comes back empty. When that happens the response carries the reasoning and a FinishReason of Length, so an empty answer is diagnosable rather than silent:

if (r.FinishReason == ChatFinishReason.Length && string.IsNullOrWhiteSpace(r.Text))
    Console.WriteLine("(no answer — the reasoning consumed the budget; raise MaxOutputTokens)");

Conversation-options template (per-client)

ChatOptions covers the per-request knobs (sampler, MaxOutputTokens, tools, thinking), but a few conversation-level settings have no MEAI surface: SystemMessage, LoraPath / AudioLoraPath, StreamToolCalls, VisualTokenBudget, FilterThinkingFromKvCache, ExtraContext, and a session-default MaxOutputTokens. Supply them once as a per-client template (LiteRtConversationOptions) on the constructor or the DI registration, and they apply to every call:

using LiteRtLmSharp;
using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;

var template = new LiteRtConversationOptions
{
    SystemMessage = "You are a terse, on-device assistant.",
    VisualTokenBudget = 256,          // cap what an image costs during prefill
    FilterThinkingFromKvCache = true, // keep long reasoning out of later turns' context
};

using IChatClient client = new LiteRtChatClient(engine, modelId: "gemma-4-E2B-it", optionsTemplate: template);

// ... or via DI (the template flows to the registered client):
services.AddLiteRtChatClient(engineOptions, optionsTemplate: template);

The merge is per-call-wins: any value the request's ChatOptions supplies (sampler, thinking, constrained decoding, tools) overrides the template, and the template fills the rest. The SystemMessage rule is the one to note: a system message on the request (the leading system chat message) always wins, and the template's SystemMessage is used only when the request carries none, so there are never two system turns. The template must not set History or HistoryJson (history is always per-call); doing so throws at construction (or at registration, for the DI overloads).

When the template sets StreamToolCalls = true, the streaming path surfaces each raw tool-call fragment as a content-less ChatResponseUpdate whose AdditionalProperties carries the fragment under the key litertlm.tool_call_delta. Use it for progress display only, and act on the complete FunctionCallContent that the following tool-call update carries. Without StreamToolCalls no such updates are emitted, so it is invisible unless you opt in.

Stateful conversations (opt-in)

By default the client is stateless: it rebuilds a fresh LiteRtConversation from the full message list every call (see Design). You can instead opt into stateful conversations, where the client keeps the live native conversation alive between calls and resumes it, so each turn re-prefills only the new messages rather than the whole thread. This is Microsoft.Extensions.AI's canonical stateful-provider contract, built on ChatResponse.ConversationId / ChatOptions.ConversationId.

Pass a LiteRtStatefulConversationOptions to the constructor (or a DI registration) to turn it on:

using LiteRtLmSharp;
using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;

using IChatClient client = new LiteRtChatClient(
    engine, modelId: "gemma-4-E2B-it",
    statefulConversations: new LiteRtStatefulConversationOptions());

// ... or via DI (the option flows to the registered client):
services.AddLiteRtChatClient(engineOptions, statefulConversations: new LiteRtStatefulConversationOptions());

Then follow the canonical MEAI loop: the first response carries a ConversationId; set it on the options and send only the new user message on each subsequent turn.

var options = new ChatOptions();
List<ChatMessage> outgoing = [new(ChatRole.User, "Remember that my favorite color is teal.")];

while (true)
{
    ChatResponse response = await client.GetResponseAsync(outgoing, options);
    Console.WriteLine(response.Text);

    // Reuse the id and send ONLY the next turn's new messages (the live conversation keeps the rest):
    options.ConversationId = response.ConversationId;
    outgoing = [new(ChatRole.User, Console.ReadLine()!)];
}

Because the response now carries a non-null ConversationId, UseFunctionInvocation() becomes incremental too: FunctionInvokingChatClient sends only the new function-result message(s) on its next iteration (with the id set) instead of replaying the accumulated history, so multi-round tool loops resume the same live conversation with no extra work on your part.

What is fixed once a conversation is created (and therefore ignored on a continuation, i.e. any call that carries a ConversationId):

  • the sampler, thinking mode, tools, constrained decoding, the system message, and any template values;
  • options that map to native per-send settings still apply on a continuation: MaxOutputTokens, FrequencyPenalty/PresencePenalty, and a JSON-schema ResponseFormat (see the structured-output note below — the conversation must have carried a schema on its first call, or the client's options template must set ConstraintProvider; otherwise the continuation throws ArgumentException and the conversation stays resumable);
  • a system message on a continuation throws InvalidOperationException (the preface cannot be rewritten). To change any fixed setting, start a new conversation by omitting ConversationId.

Lifetime and eviction:

  • live conversations are held in an LRU cache bounded by MaxLiveConversations (default 8);
  • creating a new conversation beyond the cap evicts and disposes the least-recently-used one;
  • a request whose ConversationId was evicted, or was never issued by this client, throws ArgumentException;
  • disposing the client disposes every live conversation. There is no time-based expiry in this mode, so size the cap for your concurrency.

Multiple live conversations require native LiteRT-LM v0.15.0+. Earlier runtimes silently lost a suspended conversation's state whenever another conversation advanced, so this mode was hard-limited to a single live conversation (and forking was unavailable) until that fix shipped.

Forking a conversation

In the stateful mode the client also offers a provider-specific forking hatch: branch a live conversation into an independent copy that shares the parent's prefilled context (a cheap native KV-cache clone — no re-prefill of the shared prefix) and diverges from there. Useful for exploring several continuations of one prompt (best-of-N, A/B prompting, tree search) without re-establishing the common prefix each time.

var branching = (LiteRtConversationBranching)client.GetService(typeof(LiteRtConversationBranching))!;

ChatResponse seeded = await client.GetResponseAsync(seedMessages, options);
string branchId = await branching.ForkAsync(seeded.ConversationId!);

// The branch id behaves like any other ConversationId: resume it, re-fork it, or let it be evicted.
options.ConversationId = branchId;

GetService(typeof(LiteRtConversationBranching)) returns the hatch only in the stateful mode (null otherwise — there are no live conversations to fork). The fork counts toward MaxLiveConversations like any other live conversation, and it inherits the parent's synthesized-call-id map, so function-calling continuations keep resolving on the branch.

ChatResponse.Usage.TotalTokenCount in this mode is the conversation's cumulative KV-cache size, so it grows across the thread (rather than resetting per call as in the stateless mode).

The context limit is a hard wall — let the guard police it. A long stateful thread (a multi-round tool loop especially) eventually approaches the engine's MaxNumTokens, and the native runtime does not check it: an unguarded overflow corrupts native memory and crashes the process on a later call. Load the engine with an explicit MaxNumTokens (as the snippets above do) to arm the binding's KV overflow guard: replies are clamped to the remaining context, and a send that no longer fits throws LiteRtContextOverflowException instead (sends carrying media are the exception — their prefill cost is not measurable managed-side, so they get only the conversation-full check; leave headroom when attachments are in play). What happens to the live conversation depends on which rejection you got: a "message doesn't fit" rejection happens before any native work, so the conversation survives — retry the same ConversationId with a shorter message. A "context is full" rejection is terminal: the client evicts the live conversation (resuming its id throws ArgumentException, like any evicted id) — start a new conversation, carrying over a summary or a trimmed history if the thread must continue.

You do not have to wait for that exception to find out: a reply that filled the context carries ChatFinishReason.Length in the same turn (on the blocking response, and on the final update of a clamped stream — whose only other symptom is that it just stops). Treat Length as "this conversation is over": even when the response carries tool calls, sending their results would throw, which is why Length deliberately wins over ToolCalls there. Note that FunctionInvokingChatClient loops on function-call content regardless of finish reason, so an unattended tool loop still terminates in the exception; the signal is for callers who look. To act before hitting the wall, track Usage.TotalTokenCount against MaxNumTokens and wind down early.

Multimodal (image / audio)

On a multimodal model, attach an image or audio clip to the final user message as a DataContent (inline bytes) or a file-path UriContent, with an image/* or audio/* media type:

using Microsoft.Extensions.AI;

byte[] png = File.ReadAllBytes("photo.png");
var message = new ChatMessage(ChatRole.User,
[
    new TextContent("What is in this image?"),
    new DataContent(png, "image/png"),
]);

ChatResponse response = await client.GetResponseAsync([message]);

The engine must have been loaded with the matching modality enabled — LiteRtEngineOptions.VisionBackend for images, AudioBackend for audio — on a multimodal model (e.g. the Gemma 4 E-series). Without it the send throws with a message naming the likely cause.

Only the final (triggering) user message's media is sent; media on earlier history turns is not replayed (the stateless connector restores prior turns as text). Remote (non-file://) URIs are skipped — the on-device engine cannot fetch them, so supply bytes or a local file.

Token usage

Every response carries ChatResponse.Usage. TotalTokenCount — the turn's prompt + reply, read from the conversation's KV cache — is always set, at no cost:

ChatResponse response = await client.GetResponseAsync("Hello", options);
long? total = response.Usage?.TotalTokenCount;   // e.g. to track how full the context window is

The input/output split (InputTokenCount / OutputTokenCount) is populated only when the engine was loaded with EnableBenchmark = true — it comes from the engine's benchmark counters (the overhead is just timing bookkeeping). Without it those stay null, and a note is left under response.AdditionalProperties (key litertlm.usage_note) explaining how to enable them:

using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
    ModelPath = "…", Backend = LiteRtBackend.Cpu, EnableBenchmark = true,   // also enables Input/OutputTokenCount
});
// …
long? input = response.Usage?.InputTokenCount;     // prefill tokens
long? output = response.Usage?.OutputTokenCount;   // decode tokens

When streaming, the usage arrives as a final UsageContent update, which MEAI aggregates into the response's Usage.

Design

  • Stateless by default. IChatClient hands the full message list every call, so the client rebuilds a fresh LiteRtConversation each time — prior messages restored as History (replayed through prefill), the final user turn sent. This keeps the caller's history and the model's KV cache in lockstep. The cost is an O(history) prefill per turn; for very long chats, opt into stateful conversations (keep the live native conversation alive between calls and re-prefill only the new turn) or drive the native LiteRtConversation API directly.
  • Serialized. LiteRtLmSharp allows one live engine per process and conversations are not thread-safe, so the client serializes calls through an internal SemaphoreSlim.
  • Engine ownership. Pass a LiteRtEngine you own (you dispose it), or register from LiteRtEngineOptions so the container loads/owns/disposes a single shared engine. The same AddLiteRtChatClient(options) registration makes the IChatClient available to MAF, Semantic Kernel and plain MEAI at once.
  • Message roles. system / user / assistant / tool are handled. The list must end with a user message, or a tool message (the function-calling continuation, appended by UseFunctionInvocation() / Semantic Kernel); the assistant tool-call turn is restored as history and the tool results are returned. A system message is restored through the History path (as a leading LiteRtMessage.System(...)), so this connector was never affected by the pre-v0.14.0 LiteRtConversationOptions.SystemMessage bug; it does not use that property.
  • Engine options. Any LiteRtEngineOptions you pass to AddLiteRtChatClient / new LiteRtChatClient flows straight through, including the ones added in v0.14.0 (NumThreads / AudioNumThreads, the LoRA ranks); the connector surface is unchanged.
  • Testing your app code. IChatClient is the intended seam for unit tests: have your code depend on IChatClient and substitute a mock/stub in tests — no model file or native binaries needed. The core types (LiteRtEngine / LiteRtConversation) are sealed and bound to the native runtime; code that drives them directly is best covered by wrapping them behind your own abstraction, or by integration tests against a real model.

Scope

  • Tool calling is supported (see Function calling).
  • Embeddings: the LiteRT-LM C API exposes no embeddings functions at v0.14.0, so there is no IEmbeddingGenerator.
  • Not AOT/trim-clean. The core LiteRtLmSharp package stays AOT/trim-friendly; this companion does not carry that guarantee.

Sample

The Semantic Kernel console sample in samples/SemanticKernel also resolves the underlying IChatClient from the kernel — see docs/semantic-kernel.md for the Semantic Kernel layer.