Semantic Kernel integration
LiteRtLmSharp.SemanticKernel is a separate, optional companion package that plugs a LiteRtLmSharp
on-device model into Microsoft Semantic Kernel as a
standard IChatCompletionService.
It is a thin layer over the Microsoft.Extensions.AI IChatClient: it registers the
LiteRtLmSharp IChatClient and exposes it to Semantic Kernel through SK's own
AsChatCompletionService
adapter. So all of Semantic Kernel's chat machinery — message conversion and function calling — flows
through that one chat client, and the underlying model is simultaneously available to
Microsoft Agent Framework and plain MEAI from the same registration.
| Package | Depends on | Target |
|---|---|---|
LiteRtLmSharp.SemanticKernel |
LiteRtLmSharp + LiteRtLmSharp.Extensions.AI (same version) + Microsoft.SemanticKernel.Abstractions 1.77.0 + Microsoft.Extensions.AI 10.7.0 |
net10.0 |
<PackageReference Include="LiteRtLmSharp" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.SemanticKernel" Version="1.1.1" />
Quick start
using LiteRtLmSharp;
using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.ChatCompletion;
// One call is the whole integration: the container loads/owns/disposes a single shared engine and exposes
// the model as an IChatCompletionService (and, underneath, as a Microsoft.Extensions.AI IChatClient).
var builder = Kernel.CreateBuilder();
builder.AddLiteRtChatCompletion(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm",
Backend = LiteRtBackend.Cpu,
MaxNumTokens = 4096,
});
Kernel kernel = builder.Build();
// From here it is ordinary Semantic Kernel.
Console.WriteLine(await kernel.InvokePromptAsync("Write one upbeat sentence about on-device AI."));
eager: true loads the weights at registration (surfacing a bad model path/backend there) instead of lazily
on first use. To use an engine you already loaded (and will dispose yourself), pass it instead of options:
builder.AddLiteRtChatCompletion(engine, modelId: "gemma-4-E2B-it"). Both overloads exist on IKernelBuilder
and IServiceCollection, and take an optional serviceId for keyed registration (e.g. an on-device service
alongside a cloud one).
Chat with history (streaming)
IChatCompletionService chat = kernel.GetRequiredService<IChatCompletionService>();
var settings = new LiteRtPromptExecutionSettings { Temperature = 0.8f, MaxTokens = 256 };
var history = new ChatHistory("You are a concise, friendly assistant.");
history.AddUserMessage("Hi! What are you good at?");
var reply = new StringBuilder();
await foreach (var chunk in chat.GetStreamingChatMessageContentsAsync(history, settings, kernel))
{
Console.Write(chunk.Content);
reply.Append(chunk.Content);
}
history.AddAssistantMessage(reply.ToString()); // record the turn so the next call sees it
Execution settings
LiteRtPromptExecutionSettings carries the sampler/output knobs. Every property is optional; unset values
fall back to the engine/model defaults. The values are stored in ExtensionData under the well-known keys
below, so they flow through Semantic Kernel's PromptExecutionSettings → ChatOptions conversion to the
chat client — and a plain PromptExecutionSettings whose ExtensionData holds the same keys (e.g. from a
prompt template's YAML) works equally well.
| Property | Key | Maps to |
|---|---|---|
Temperature |
temperature |
sampler temperature |
TopP |
top_p |
sampler top-p (selects the TopP sampler) |
TopK |
top_k |
sampler top-k |
MaxTokens |
max_tokens |
max output tokens |
Seed |
seed |
sampler seed |
EnableThinking |
enable_thinking |
reasoning mode (see below) |
EnableConstrainedDecoding |
enable_constrained_decoding |
force schema-constrained tool-call arguments (Function calling) |
NoRepeatNgramSize |
no_repeat_ngram_size |
ban repeating any n-gram of that size within the reply (native v0.15.0+) |
SuppressTokens |
suppress_tokens |
token ids that can never be sampled — an int[] (a JSON array works from prompt templates); find ids with LiteRtEngine.Tokenize (native v0.15.0+) |
Conversation-options template (per-service)
LiteRtPromptExecutionSettings covers the per-request sampler/output knobs, but a few conversation-level
settings have no execution-settings surface: SystemMessage, LoraPath / AudioLoraPath, StreamToolCalls,
VisualTokenBudget, FilterThinkingFromKvCache, ExtraContext, and a session-default MaxOutputTokens.
Pass them once as a LiteRtConversationOptions template on the registration and they apply to every call:
using LiteRtLmSharp;
using Microsoft.SemanticKernel;
var template = new LiteRtConversationOptions
{
SystemMessage = "You are a terse, on-device assistant.",
VisualTokenBudget = 256,
};
var builder = Kernel.CreateBuilder();
builder.AddLiteRtChatCompletion(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", Backend = LiteRtBackend.Cpu, MaxNumTokens = 4096,
}, optionsTemplate: template);
The merge is per-call-wins: anything the request's execution settings supply (sampler, thinking, constrained
decoding, function choice) overrides the template, and the template fills the rest. A ChatHistory system
message on the request always wins, and the template's SystemMessage is used only when the request carries
none, so there are never two system turns. The template must not carry History / HistoryJson (history is
per-call); doing so throws at registration. The parameter is available on every AddLiteRtChatCompletion
overload, on both IKernelBuilder and IServiceCollection.
Design: a stateless connector over a stateful engine
Semantic Kernel's IChatCompletionService (and the IChatClient underneath) is stateless: the caller
passes the full ChatHistory on every call. A LiteRtConversation, by contrast, is stateful — it holds
a KV cache that accumulates across turns.
The connector bridges the two by being stateless on every call: it builds a fresh conversation each
time, restoring all-but-last messages as History (replayed through prefill) and
sending the final user turn to trigger generation.
ChatHistory: [System, User₁, Assistant₁, User₂] (what SK hands the connector)
└──────── History ────────┘ └ Send ┘
(re-prefilled each call) (generates)
This keeps SK's history and the model's KV cache in lockstep (SK owns the history and can edit it). The cost
is an O(history) prefill per turn — fine for typical chats; for very long conversations, drive the native
LiteRtConversation API directly. Calls are serialized (one live engine per
process; conversations are not thread-safe), and the engine lifecycle is handled by the
Extensions.AI registration the connector builds on. A ChatHistory system message is
restored through that same History path, so this connector was never affected by the pre-v0.14.0
LiteRtConversationOptions.SystemMessage bug. And any LiteRtEngineOptions you pass to
AddLiteRtChatCompletion, including the v0.14.0 additions (NumThreads, LoRA ranks), flows straight
through.
Function calling
Set a FunctionChoiceBehavior on the execution settings and Semantic Kernel's functions are offered to the
model; when the model calls one, it is auto-invoked and the result fed back so the model answers from it —
the standard SK function-calling experience.
using System.ComponentModel;
using Microsoft.SemanticKernel;
// A plugin the model may call.
sealed class WeatherPlugin
{
[KernelFunction, Description("Gets the current weather for a city.")]
public string GetWeather([Description("The city.")] string city) => $"22°C and sunny in {city}";
}
kernel.Plugins.AddFromObject(new WeatherPlugin());
var settings = new LiteRtPromptExecutionSettings
{
FunctionChoiceBehavior = FunctionChoiceBehavior.Auto(),
EnableConstrainedDecoding = true, // recommended for small models (see note); off by default
};
var answer = await kernel.InvokePromptAsync(
"What's the weather in Paris? Use the weather tool.", new KernelArguments(settings));
Console.WriteLine(answer); // "The weather in Paris is 22°C and sunny."
How it works. Semantic Kernel's AsChatCompletionService adapter passes the kernel's functions to the
chat client as tools but does not run the auto-invoke loop itself, so the connector wraps the IChatClient
with MEAI's function-invocation middleware (UseFunctionInvocation). That wrapping is built into
AddLiteRtChatCompletion, so FunctionChoiceBehavior.Auto() just works; it is a no-op when a request carries
no functions. Under the hood this is the same bridge the Microsoft.Extensions.AI integration
exposes (the model's tool calls become FunctionCallContent).
Constrained decoding. LiteRtPromptExecutionSettings.EnableConstrainedDecoding = true makes the model emit
valid, schema-shaped tool-call arguments — recommended for small on-device models. It is off by default and
works on every platform (the linux-x64 restriction of earlier releases is gone with the official LiteRT-LM
v0.16.0 prebuilts, which embed the constraint provider). Tools still work without it; arguments are just not
grammar-constrained.
Function choice (FunctionChoiceBehavior). Auto() (the model decides), None() (no functions offered)
and Required() (the model must call one) flow through to the chat client's ChatOptions.ToolMode. Required
is best-effort on-device — the connector instructs the model to call a function and narrows the offered set,
but cannot force a call the way a cloud API's server-enforced tool_choice: "required" does. See
Tool choice for the per-mode behavior.
Structured output (response_format)
A raw JSON-Schema JsonElement under the standard response_format execution-settings key is
enforced during sampling (native LiteRT-LM v0.15.0+): SK's adapter converts it into a
Microsoft.Extensions.AI schema ResponseFormat, and the connector arms the LlGuidance constraint
provider so the reply is guaranteed to be a conforming JSON document.
using JsonDocument schema = JsonDocument.Parse(
"""{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}""");
var settings = new LiteRtPromptExecutionSettings
{
MaxTokens = 64,
ExtensionData = new Dictionary<string, object> { ["response_format"] = schema.RootElement },
};
// result[0].Content parses as JSON conforming to the schema — enforced, not prompted.
Supported shapes (verified against a real model): the raw schema element shown above works; an
OpenAI-style {"type":"json_schema","json_schema":{...}} envelope does not (SK forwards the
whole envelope as the schema and the constraint compiler rejects it — the send fails); a plain
"json_object" string is JSON mode, which is not enforced (prompt-driven only). The same rules
and caveats as the MEAI connector apply — most importantly, not combinable with function calling
on the same request (see the structured output section
of the Extensions.AI guide).
Reasoning (thinking) and the output-token budget
With EnableThinking = true the model emits a reasoning trace before the answer, and that trace shares
the MaxTokens output budget with the answer. So give thinking models headroom: with too small a budget
the reasoning can consume it all and the answer comes back empty. For example, on gemma-4 the reasoning
for a simple prompt runs ~200 tokens, so MaxTokens = 200 with thinking on leaves nothing for the answer,
whereas MaxTokens = 512 works. This affects both blocking and streaming equally — it is a budget effect,
not a connector bug.
A caveat specific to the Semantic Kernel path: SK's IChatCompletionService adapter does not surface the
reasoning trace or a truncation signal — kernel.InvokePromptAsync / GetChatMessageContentsAsync return
just the answer text (empty when the reasoning ate the budget). (Token usage, by contrast, does flow through —
SK exposes it on the result's metadata; see Token usage.) The underlying
Microsoft.Extensions.AI IChatClient does surface both: the reasoning as a
TextReasoningContent (excluded from ChatResponse.Text), and a ChatResponse.FinishReason of Length
when the answer was truncated. The same registration provides it, so resolve it from the kernel when you need
rich reasoning / truncation handling:
using LiteRtLmSharp.Extensions.AI; // for LiteRtChatOptions
using Microsoft.Extensions.AI;
using Microsoft.Extensions.DependencyInjection;
IChatClient chatClient = kernel.Services.GetRequiredService<IChatClient>();
ChatResponse response = await chatClient.GetResponseAsync(messages,
new LiteRtChatOptions { MaxOutputTokens = 512, EnableThinking = true });
string? reasoning = response.Messages.SelectMany(m => m.Contents).OfType<TextReasoningContent>().FirstOrDefault()?.Text;
if (response.FinishReason == ChatFinishReason.Length && string.IsNullOrWhiteSpace(response.Text))
Console.WriteLine("(no answer — the reasoning consumed the budget; raise MaxOutputTokens)");
Multimodal (image / audio)
On a multimodal model, add an ImageContent or AudioContent (with inline data) to a chat message — Semantic
Kernel's adapter forwards it to the model as an attachment through the underlying IChatClient:
using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.ChatCompletion;
byte[] png = File.ReadAllBytes("photo.png");
var history = new ChatHistory();
history.Add(new ChatMessageContent(AuthorRole.User, new ChatMessageContentItemCollection
{
new TextContent("What is in this image?"),
new ImageContent(png, "image/png"),
}));
IChatCompletionService chat = kernel.GetRequiredService<IChatCompletionService>();
var reply = await chat.GetChatMessageContentsAsync(history);
The engine must have been loaded with the matching modality enabled (LiteRtEngineOptions.VisionBackend /
AudioBackend) on a multimodal model. Only the final user message's media is sent (prior turns are restored as
text). See docs/extensions-ai.md for the details.
Scope
- Function calling is supported (see Function calling) via
FunctionChoiceBehavior. - Multimodal (image / audio) is supported (see Multimodal).
- Text generation (
ITextGenerationService) is not provided — Semantic Kernel and the wider .NET AI stack are chat-centric; use chat completion. - Embeddings are not provided (the LiteRT-LM C API exposes none at v0.14.0).
- Not AOT/trim-clean. Semantic Kernel itself is not; the core
LiteRtLmSharppackage stays AOT/trim-friendly, this companion does not carry that guarantee.
Sample
A runnable console sample is in samples/SemanticKernel: it builds a kernel with
AddLiteRtChatCompletion and demonstrates a prompt function, a streaming prompt, a multi-turn streaming
chat, and function calling (a [KernelFunction] plugin) — pass --interactive for a chat loop. See its README for how to
run it. The broader Microsoft.Extensions.AI / Agent Framework story is in docs/extensions-ai.md.