Microsoft.Extensions.AI integration (IChatClient)
LiteRtLmSharp.Extensions.AI is a separate, optional companion package that exposes a LiteRtLmSharp
on-device model as a Microsoft.Extensions.AI.IChatClient.
IChatClient is the .NET ecosystem's provider-agnostic abstraction for chat models, so this one package
makes the model usable from:
- Microsoft Agent Framework (MAF) —
new ChatClientAgent(chatClient, …); - Semantic Kernel — via the
LiteRtLmSharp.SemanticKernelconnector (a thin layer over thisIChatClient); - plain Microsoft.Extensions.AI — the middleware pipeline (function invocation, caching, telemetry), dependency injection, ASP.NET Core, etc.
This is the foundation: it depends only on lightweight abstractions
(Microsoft.Extensions.AI.Abstractions + Microsoft.Extensions.DependencyInjection.Abstractions for the DI
helpers), and the Semantic Kernel connector is built on top of it.
| Package | Depends on | Target |
|---|---|---|
LiteRtLmSharp.Extensions.AI |
LiteRtLmSharp (same version) + Microsoft.Extensions.AI.Abstractions 10.7.0 + Microsoft.Extensions.DependencyInjection.Abstractions 10.0.9 |
net10.0 |
<PackageReference Include="LiteRtLmSharp" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.Extensions.AI" Version="1.1.1" />
Quick start
using LiteRtLmSharp;
using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm",
Backend = LiteRtBackend.Cpu,
MaxNumTokens = 4096,
});
using IChatClient client = new LiteRtChatClient(engine, modelId: "gemma-4-E2B-it");
// Blocking:
ChatResponse response = await client.GetResponseAsync("Write one upbeat sentence about on-device AI.");
Console.WriteLine(response.Text);
// Streaming:
await foreach (ChatResponseUpdate update in client.GetStreamingResponseAsync("Tell me a joke"))
Console.Write(update.Text);
Dependency injection
// Container loads, owns and disposes a single shared engine (one engine per process):
services.AddLiteRtChatClient(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", Backend = LiteRtBackend.Cpu, MaxNumTokens = 4096,
});
// ... or register over an engine you already loaded (you dispose it):
services.AddLiteRtChatClient(engine, modelId: "gemma-4-E2B-it");
// Elsewhere:
IChatClient client = serviceProvider.GetRequiredService<IChatClient>();
eager: true on the options overload loads the weights at registration (so a bad model path/backend throws
there) instead of lazily on first use.
Use with Microsoft Agent Framework
MAF builds agents over any IChatClient — there is no separate "MAF connector" to write:
using Microsoft.Agents.AI;
var agent = new ChatClientAgent(client, instructions: "You are a concise, helpful assistant.");
AgentRunResponse reply = await agent.RunAsync("What can you do?");
Use with the Microsoft.Extensions.AI pipeline
Because it's a standard IChatClient, it composes with the ecosystem's middleware (from the
Microsoft.Extensions.AI package) — caching, telemetry, logging, and function invocation:
IChatClient pipeline = client.AsBuilder()
.UseFunctionInvocation() // auto-invokes AIFunctions returned in ChatOptions.Tools
// .UseDistributedCache(...) // caching, telemetry, logging, … all compose here
.Build();
Function calling (tools)
Function tools in ChatOptions.Tools are passed to the model; the model's tool calls are surfaced as
FunctionCallContent (with ChatFinishReason.ToolCalls). Compose UseFunctionInvocation() and the pipeline
runs the loop: the model emits a call → the matching AIFunction is invoked → its result is returned to the
model → the model answers from it.
using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;
AIFunction getWeather = AIFunctionFactory.Create(
(string city) => $"22°C and sunny in {city}",
name: "get_weather", description: "Gets the current weather for a city.");
IChatClient client = new LiteRtChatClient(engine)
.AsBuilder()
.UseFunctionInvocation() // drives the tool loop
.Build();
var options = new LiteRtChatOptions
{
Tools = [getWeather],
EnableConstrainedDecoding = true, // recommended (see below); off by default
};
ChatResponse response = await client.GetResponseAsync("What's the weather in Paris?", options);
Console.WriteLine(response.Text); // "The weather in Paris is 22°C and sunny."
Constrained decoding. Set LiteRtChatOptions.EnableConstrainedDecoding = true so the model emits valid,
schema-shaped tool-call arguments — strongly recommended for small on-device models. It is off by default and
works on every platform (the linux-x64 restriction of earlier releases is gone: the official LiteRT-LM v0.16.0
prebuilts embed the constraint provider). Tools still work without it; arguments are just not
grammar-constrained. On a plain ChatOptions, set the enable_constrained_decoding key in
AdditionalProperties instead.
Tool choice (ChatOptions.ToolMode). The native API has no tool_choice parameter (the model always
decides), so the connector emulates the MEAI modes as best it can:
ToolMode |
Behavior |
|---|---|
Auto (default) |
All tools are offered; the model decides whether to call one. |
None |
No tools are offered, so the model can't call any. |
RequireAny |
All tools offered + a system-prompt instruction to call one — best-effort. |
RequireSpecific(name) |
Only that tool is offered, plus the instruction naming it — best-effort. |
RequireAny / RequireSpecific are best-effort, not a guarantee: unlike a cloud API's server-enforced
tool_choice: "required", the connector can only instruct the on-device model (and narrow the tool list) —
the decoder isn't forced. In practice a capable model calls the tool when the request is plausibly related to
it, but may ignore the instruction for a clearly-unrelated prompt. (Semantic Kernel's
FunctionChoiceBehavior.Auto()/None()/Required() map to these same modes.)
The native LiteRtLmSharp tools API is still available for full control (constrained decoding, custom tool-call parsing); the chat client is the MEAI-idiomatic path on top of it.
Structured output (ResponseFormat)
A JSON-schema ChatOptions.ResponseFormat is enforced during sampling (native LiteRT-LM v0.15.0+):
the client arms the LlGuidance constraint provider on the conversation and attaches the schema as a
per-send constraint, so every token the model emits is masked to schema-conforming continuations — the
reply is guaranteed to be a JSON document matching the schema, not just nudged toward it.
ChatResponse r = await client.GetResponseAsync(messages, new ChatOptions
{
ResponseFormat = ChatResponseFormat.ForJsonSchema(mySchemaJsonElement),
});
// r.Text parses as JSON conforming to mySchemaJsonElement — enforced, not prompted.
Rules and caveats:
- Schema-less
ChatResponseFormat.Json("JSON mode") is not enforced — without a schema there is nothing precise to constrain, so it keeps the previous prompt-driven behavior. - Not combinable with
Toolson one request. The schema masks every generated token, which makes emitting a tool call impossible, so a request carrying both throwsArgumentException. Run the tool phase first, then request the schema-formatted answer in a separate call. In stateful mode both phases can share one conversation: setConstraintProvideron the client's conversation-options template, run the tool request without aResponseFormat, then send the schema-formatted request as a continuation of the sameConversationId— the reply is schema-enforced and informed by the tool results (covered by a model test). - Stateful mode: the schema must be present on the conversation's first call (the constraint
provider is fixed at creation). A schema arriving only on a continuation throws
ArgumentExceptionand leaves the conversation resumable; to make every conversation schema-capable, setLiteRtConversationOptions.ConstraintProvider = LiteRtConstraintProvider.LlGuidanceon the client's conversation-options template. - When a schema request is active, the tool-calling
enable_constrained_decodingknob is dropped for that conversation (the two constrained-decoding modes are mutually exclusive natively, and the tool path is meaningless without tools).
Banning repeats and tokens: NoRepeatNgramSize / SuppressTokens
ChatOptions.FrequencyPenalty / PresencePenalty already map to the native repetition penalties. Two more
logit processors have no MEAI property, so LiteRtChatOptions exposes them, backed by the
no_repeat_ngram_size and suppress_tokens keys in AdditionalProperties (a plain ChatOptions with those
keys works too, including options deserialized from JSON/YAML):
int[] banned = [.. engine.Tokenize(" Paris"), .. engine.Tokenize("Paris")]; // both surface forms
var options = new LiteRtChatOptions
{
MaxOutputTokens = 128,
NoRepeatNgramSize = 3, // never repeat a 3-gram already produced in this reply
SuppressTokens = banned, // these ids can never be sampled
};
NoRepeatNgramSizebans any n-gram of that many tokens the reply already produced (whole-reply window). Useful when a small model echoes a prompt template verbatim. Zero or negative is rejected.SuppressTokensforces the listed ids' logits to-infon every step. Find ids withLiteRtEngine.Tokenize; most words tokenize differently with and without a leading space, so ban both. Negative ids are rejected; ids outside the vocabulary are ignored natively.- Both apply per request and require native LiteRT-LM v0.15.0+. The Semantic Kernel connector exposes the
same two knobs on
LiteRtPromptExecutionSettings.
Reasoning ("thinking")
Enable the model's reasoning mode with LiteRtChatOptions — a ChatOptions subtype that adds the knobs
MEAI has no typed property for (EnableThinking, EnableConstrainedDecoding, NoRepeatNgramSize,
SuppressTokens, each backed by a key in AdditionalProperties). This mirrors the Semantic Kernel
connector's LiteRtPromptExecutionSettings:
var options = new LiteRtChatOptions { MaxOutputTokens = 512, EnableThinking = true };
ChatResponse r = await client.GetResponseAsync(messages, options);
(LiteRtChatOptions is a ChatOptions, so it works anywhere one is accepted — ChatOptions already exposes
Temperature/TopP/TopK/MaxOutputTokens/Seed. Setting the enable_thinking key directly on
ChatOptions.AdditionalProperties is equivalent.)
The reasoning trace is surfaced as a TextReasoningContent on the response (and as reasoning updates
when streaming). It is excluded from ChatResponse.Text, so the answer stays clean while the reasoning
stays accessible:
string answer = r.Text;
string? reasoning = r.Messages.SelectMany(m => m.Contents).OfType<TextReasoningContent>().FirstOrDefault()?.Text;
The reasoning shares the MaxOutputTokens budget with the answer. Give thinking models headroom — with
too small a budget the reasoning can consume it and the answer comes back empty. When that happens the
response carries the reasoning and a FinishReason of Length, so an empty answer is diagnosable rather
than silent:
if (r.FinishReason == ChatFinishReason.Length && string.IsNullOrWhiteSpace(r.Text))
Console.WriteLine("(no answer — the reasoning consumed the budget; raise MaxOutputTokens)");
Conversation-options template (per-client)
ChatOptions covers the per-request knobs (sampler, MaxOutputTokens, tools, thinking), but a few
conversation-level settings have no MEAI surface: SystemMessage, LoraPath / AudioLoraPath,
StreamToolCalls, VisualTokenBudget, FilterThinkingFromKvCache, ExtraContext, and a session-default
MaxOutputTokens. Supply them once as a per-client template (LiteRtConversationOptions) on the
constructor or the DI registration, and they apply to every call:
using LiteRtLmSharp;
using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;
var template = new LiteRtConversationOptions
{
SystemMessage = "You are a terse, on-device assistant.",
VisualTokenBudget = 256, // cap what an image costs during prefill
FilterThinkingFromKvCache = true, // keep long reasoning out of later turns' context
};
using IChatClient client = new LiteRtChatClient(engine, modelId: "gemma-4-E2B-it", optionsTemplate: template);
// ... or via DI (the template flows to the registered client):
services.AddLiteRtChatClient(engineOptions, optionsTemplate: template);
The merge is per-call-wins: any value the request's ChatOptions supplies (sampler, thinking, constrained
decoding, tools) overrides the template, and the template fills the rest. The SystemMessage rule is the
one to note: a system message on the request (the leading system chat message) always wins, and the
template's SystemMessage is used only when the request carries none, so there are never two system
turns. The template must not set History or HistoryJson (history is always per-call); doing so throws at
construction (or at registration, for the DI overloads).
When the template sets StreamToolCalls = true, the streaming path surfaces each raw tool-call fragment as a
content-less ChatResponseUpdate whose AdditionalProperties carries the fragment under the key
litertlm.tool_call_delta. Use it for progress display only, and act on the complete FunctionCallContent
that the following tool-call update carries. Without StreamToolCalls no such updates are emitted, so it is
invisible unless you opt in.
Stateful conversations (opt-in)
By default the client is stateless: it rebuilds a fresh LiteRtConversation from the full message list
every call (see Design). You can instead opt into stateful conversations, where the client
keeps the live native conversation alive between calls and resumes it, so each turn re-prefills only the new
messages rather than the whole thread. This is Microsoft.Extensions.AI's canonical stateful-provider contract,
built on ChatResponse.ConversationId / ChatOptions.ConversationId.
Pass a LiteRtStatefulConversationOptions to the constructor (or a DI registration) to turn it on:
using LiteRtLmSharp;
using LiteRtLmSharp.Extensions.AI;
using Microsoft.Extensions.AI;
using IChatClient client = new LiteRtChatClient(
engine, modelId: "gemma-4-E2B-it",
statefulConversations: new LiteRtStatefulConversationOptions());
// ... or via DI (the option flows to the registered client):
services.AddLiteRtChatClient(engineOptions, statefulConversations: new LiteRtStatefulConversationOptions());
Then follow the canonical MEAI loop: the first response carries a ConversationId; set it on the options and
send only the new user message on each subsequent turn.
var options = new ChatOptions();
List<ChatMessage> outgoing = [new(ChatRole.User, "Remember that my favorite color is teal.")];
while (true)
{
ChatResponse response = await client.GetResponseAsync(outgoing, options);
Console.WriteLine(response.Text);
// Reuse the id and send ONLY the next turn's new messages (the live conversation keeps the rest):
options.ConversationId = response.ConversationId;
outgoing = [new(ChatRole.User, Console.ReadLine()!)];
}
Because the response now carries a non-null ConversationId, UseFunctionInvocation() becomes incremental
too: FunctionInvokingChatClient sends only the new function-result message(s) on its next iteration (with
the id set) instead of replaying the accumulated history, so multi-round tool loops resume the same live
conversation with no extra work on your part.
What is fixed once a conversation is created (and therefore ignored on a continuation, i.e. any call that
carries a ConversationId):
- the sampler, thinking mode, tools, constrained decoding, the system message, and any template values;
- options that map to native per-send settings still apply on a continuation:
MaxOutputTokens,FrequencyPenalty/PresencePenalty, and a JSON-schemaResponseFormat(see the structured-output note below — the conversation must have carried a schema on its first call, or the client's options template must setConstraintProvider; otherwise the continuation throwsArgumentExceptionand the conversation stays resumable); - a system message on a continuation throws
InvalidOperationException(the preface cannot be rewritten). To change any fixed setting, start a new conversation by omittingConversationId.
Lifetime and eviction:
- live conversations are held in an LRU cache bounded by
MaxLiveConversations(default 8); - creating a new conversation beyond the cap evicts and disposes the least-recently-used one;
- a request whose
ConversationIdwas evicted, or was never issued by this client, throwsArgumentException; - disposing the client disposes every live conversation. There is no time-based expiry in this mode, so size the cap for your concurrency.
Multiple live conversations require native LiteRT-LM v0.15.0+. Earlier runtimes silently lost a suspended conversation's state whenever another conversation advanced, so this mode was hard-limited to a single live conversation (and forking was unavailable) until that fix shipped.
Forking a conversation
In the stateful mode the client also offers a provider-specific forking hatch: branch a live conversation into an independent copy that shares the parent's prefilled context (a cheap native KV-cache clone — no re-prefill of the shared prefix) and diverges from there. Useful for exploring several continuations of one prompt (best-of-N, A/B prompting, tree search) without re-establishing the common prefix each time.
var branching = (LiteRtConversationBranching)client.GetService(typeof(LiteRtConversationBranching))!;
ChatResponse seeded = await client.GetResponseAsync(seedMessages, options);
string branchId = await branching.ForkAsync(seeded.ConversationId!);
// The branch id behaves like any other ConversationId: resume it, re-fork it, or let it be evicted.
options.ConversationId = branchId;
GetService(typeof(LiteRtConversationBranching)) returns the hatch only in the stateful mode (null
otherwise — there are no live conversations to fork). The fork counts toward MaxLiveConversations like any
other live conversation, and it inherits the parent's synthesized-call-id map, so function-calling
continuations keep resolving on the branch.
ChatResponse.Usage.TotalTokenCount in this mode is the conversation's cumulative KV-cache size, so it
grows across the thread (rather than resetting per call as in the stateless mode).
The context limit is a hard wall — let the guard police it. A long stateful thread (a multi-round tool
loop especially) eventually approaches the engine's MaxNumTokens, and the native runtime does not check
it: an unguarded overflow corrupts native memory and crashes the process on a later call. Load the engine
with an explicit MaxNumTokens (as the snippets above do) to arm the binding's KV overflow guard:
replies are clamped to the remaining context, and a send that no longer fits throws
LiteRtContextOverflowException instead (sends carrying media are the exception — their prefill cost is
not measurable managed-side, so they get only the conversation-full check; leave headroom when
attachments are in play). What happens to the live conversation depends on which rejection you got: a
"message doesn't fit" rejection happens before any native work, so the conversation survives — retry
the same ConversationId with a shorter message. A "context is full" rejection is terminal: the
client evicts the live conversation (resuming its id throws ArgumentException, like any evicted id) —
start a new conversation, carrying over a summary or a trimmed history if the thread must continue.
You do not have to wait for that exception to find out: a reply that filled the context carries
ChatFinishReason.Length in the same turn (on the blocking response, and on the final update of a
clamped stream — whose only other symptom is that it just stops). Treat Length as "this conversation is
over": even when the response carries tool calls, sending their results would throw, which is why Length
deliberately wins over ToolCalls there. Note that FunctionInvokingChatClient loops on function-call
content regardless of finish reason, so an unattended tool loop still terminates in the exception; the
signal is for callers who look. To act before hitting the wall, track Usage.TotalTokenCount against
MaxNumTokens and wind down early.
Multimodal (image / audio)
On a multimodal model, attach an image or audio clip to the final user message as a DataContent (inline
bytes) or a file-path UriContent, with an image/* or audio/* media type:
using Microsoft.Extensions.AI;
byte[] png = File.ReadAllBytes("photo.png");
var message = new ChatMessage(ChatRole.User,
[
new TextContent("What is in this image?"),
new DataContent(png, "image/png"),
]);
ChatResponse response = await client.GetResponseAsync([message]);
The engine must have been loaded with the matching modality enabled — LiteRtEngineOptions.VisionBackend for
images, AudioBackend for audio — on a multimodal model (e.g. the Gemma 4 E-series). Without it the send
throws with a message naming the likely cause.
Only the final (triggering) user message's media is sent; media on earlier history turns is not replayed
(the stateless connector restores prior turns as text). Remote (non-file://) URIs are skipped — the
on-device engine cannot fetch them, so supply bytes or a local file.
Token usage
Every response carries ChatResponse.Usage. TotalTokenCount — the turn's prompt + reply, read from the
conversation's KV cache — is always set, at no cost:
ChatResponse response = await client.GetResponseAsync("Hello", options);
long? total = response.Usage?.TotalTokenCount; // e.g. to track how full the context window is
The input/output split (InputTokenCount / OutputTokenCount) is populated only when the engine was loaded
with EnableBenchmark = true — it comes from the engine's benchmark counters (the overhead is just timing
bookkeeping). Without it those stay null, and a note is left under response.AdditionalProperties (key
litertlm.usage_note) explaining how to enable them:
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "…", Backend = LiteRtBackend.Cpu, EnableBenchmark = true, // also enables Input/OutputTokenCount
});
// …
long? input = response.Usage?.InputTokenCount; // prefill tokens
long? output = response.Usage?.OutputTokenCount; // decode tokens
When streaming, the usage arrives as a final UsageContent update, which MEAI aggregates into the response's Usage.
Design
- Stateless by default.
IChatClienthands the full message list every call, so the client rebuilds a freshLiteRtConversationeach time — prior messages restored asHistory(replayed through prefill), the final user turn sent. This keeps the caller's history and the model's KV cache in lockstep. The cost is anO(history)prefill per turn; for very long chats, opt into stateful conversations (keep the live native conversation alive between calls and re-prefill only the new turn) or drive the nativeLiteRtConversationAPI directly. - Serialized. LiteRtLmSharp allows one live engine per process and conversations are not thread-safe, so
the client serializes calls through an internal
SemaphoreSlim. - Engine ownership. Pass a
LiteRtEngineyou own (you dispose it), or register fromLiteRtEngineOptionsso the container loads/owns/disposes a single shared engine. The sameAddLiteRtChatClient(options)registration makes theIChatClientavailable to MAF, Semantic Kernel and plain MEAI at once. - Message roles.
system/user/assistant/toolare handled. The list must end with a user message, or a tool message (the function-calling continuation, appended byUseFunctionInvocation()/ Semantic Kernel); the assistant tool-call turn is restored as history and the tool results are returned. Asystemmessage is restored through theHistorypath (as a leadingLiteRtMessage.System(...)), so this connector was never affected by the pre-v0.14.0LiteRtConversationOptions.SystemMessagebug; it does not use that property. - Engine options. Any
LiteRtEngineOptionsyou pass toAddLiteRtChatClient/new LiteRtChatClientflows straight through, including the ones added in v0.14.0 (NumThreads/AudioNumThreads, the LoRA ranks); the connector surface is unchanged. - Testing your app code.
IChatClientis the intended seam for unit tests: have your code depend onIChatClientand substitute a mock/stub in tests — no model file or native binaries needed. The core types (LiteRtEngine/LiteRtConversation) are sealed and bound to the native runtime; code that drives them directly is best covered by wrapping them behind your own abstraction, or by integration tests against a real model.
Scope
- Tool calling is supported (see Function calling).
- Embeddings: the LiteRT-LM C API exposes no embeddings functions at v0.14.0, so there is no
IEmbeddingGenerator. - Not AOT/trim-clean. The core
LiteRtLmSharppackage stays AOT/trim-friendly; this companion does not carry that guarantee.
Sample
The Semantic Kernel console sample in samples/SemanticKernel also resolves
the underlying IChatClient from the kernel — see docs/semantic-kernel.md for the
Semantic Kernel layer.