Semantic Kernel integration

LiteRtLmSharp.SemanticKernel is a separate, optional companion package that plugs a LiteRtLmSharp on-device model into Microsoft Semantic Kernel as a standard IChatCompletionService.

It is a thin layer over the Microsoft.Extensions.AI IChatClient: it registers the LiteRtLmSharp IChatClient and exposes it to Semantic Kernel through SK's own AsChatCompletionService adapter. So all of Semantic Kernel's chat machinery — message conversion and function calling — flows through that one chat client, and the underlying model is simultaneously available to Microsoft Agent Framework and plain MEAI from the same registration.

Package Depends on Target
LiteRtLmSharp.SemanticKernel LiteRtLmSharp + LiteRtLmSharp.Extensions.AI (same version) + Microsoft.SemanticKernel.Abstractions 1.77.0 + Microsoft.Extensions.AI 10.7.0 net10.0
<PackageReference Include="LiteRtLmSharp" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.SemanticKernel" Version="1.1.1" />

Quick start

using LiteRtLmSharp;
using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.ChatCompletion;

// One call is the whole integration: the container loads/owns/disposes a single shared engine and exposes
// the model as an IChatCompletionService (and, underneath, as a Microsoft.Extensions.AI IChatClient).
var builder = Kernel.CreateBuilder();
builder.AddLiteRtChatCompletion(new LiteRtEngineOptions
{
    ModelPath = "gemma-4-E2B-it.litertlm",
    Backend = LiteRtBackend.Cpu,
    MaxNumTokens = 4096,
});
Kernel kernel = builder.Build();

// From here it is ordinary Semantic Kernel.
Console.WriteLine(await kernel.InvokePromptAsync("Write one upbeat sentence about on-device AI."));

eager: true loads the weights at registration (surfacing a bad model path/backend there) instead of lazily on first use. To use an engine you already loaded (and will dispose yourself), pass it instead of options: builder.AddLiteRtChatCompletion(engine, modelId: "gemma-4-E2B-it"). Both overloads exist on IKernelBuilder and IServiceCollection, and take an optional serviceId for keyed registration (e.g. an on-device service alongside a cloud one).

Chat with history (streaming)

IChatCompletionService chat = kernel.GetRequiredService<IChatCompletionService>();
var settings = new LiteRtPromptExecutionSettings { Temperature = 0.8f, MaxTokens = 256 };

var history = new ChatHistory("You are a concise, friendly assistant.");
history.AddUserMessage("Hi! What are you good at?");

var reply = new StringBuilder();
await foreach (var chunk in chat.GetStreamingChatMessageContentsAsync(history, settings, kernel))
{
    Console.Write(chunk.Content);
    reply.Append(chunk.Content);
}
history.AddAssistantMessage(reply.ToString());   // record the turn so the next call sees it

Execution settings

LiteRtPromptExecutionSettings carries the sampler/output knobs. Every property is optional; unset values fall back to the engine/model defaults. The values are stored in ExtensionData under the well-known keys below, so they flow through Semantic Kernel's PromptExecutionSettings → ChatOptions conversion to the chat client — and a plain PromptExecutionSettings whose ExtensionData holds the same keys (e.g. from a prompt template's YAML) works equally well.

Property Key Maps to
Temperature temperature sampler temperature
TopP top_p sampler top-p (selects the TopP sampler)
TopK top_k sampler top-k
MaxTokens max_tokens max output tokens
Seed seed sampler seed
EnableThinking enable_thinking reasoning mode (see below)
EnableConstrainedDecoding enable_constrained_decoding force schema-constrained tool-call arguments (Function calling)
NoRepeatNgramSize no_repeat_ngram_size ban repeating any n-gram of that size within the reply (native v0.15.0+)
SuppressTokens suppress_tokens token ids that can never be sampled — an int[] (a JSON array works from prompt templates); find ids with LiteRtEngine.Tokenize (native v0.15.0+)

Conversation-options template (per-service)

LiteRtPromptExecutionSettings covers the per-request sampler/output knobs, but a few conversation-level settings have no execution-settings surface: SystemMessage, LoraPath / AudioLoraPath, StreamToolCalls, VisualTokenBudget, FilterThinkingFromKvCache, ExtraContext, and a session-default MaxOutputTokens. Pass them once as a LiteRtConversationOptions template on the registration and they apply to every call:

using LiteRtLmSharp;
using Microsoft.SemanticKernel;

var template = new LiteRtConversationOptions
{
    SystemMessage = "You are a terse, on-device assistant.",
    VisualTokenBudget = 256,
};

var builder = Kernel.CreateBuilder();
builder.AddLiteRtChatCompletion(new LiteRtEngineOptions
{
    ModelPath = "gemma-4-E2B-it.litertlm", Backend = LiteRtBackend.Cpu, MaxNumTokens = 4096,
}, optionsTemplate: template);

The merge is per-call-wins: anything the request's execution settings supply (sampler, thinking, constrained decoding, function choice) overrides the template, and the template fills the rest. A ChatHistory system message on the request always wins, and the template's SystemMessage is used only when the request carries none, so there are never two system turns. The template must not carry History / HistoryJson (history is per-call); doing so throws at registration. The parameter is available on every AddLiteRtChatCompletion overload, on both IKernelBuilder and IServiceCollection.

Design: a stateless connector over a stateful engine

Semantic Kernel's IChatCompletionService (and the IChatClient underneath) is stateless: the caller passes the full ChatHistory on every call. A LiteRtConversation, by contrast, is stateful — it holds a KV cache that accumulates across turns.

The connector bridges the two by being stateless on every call: it builds a fresh conversation each time, restoring all-but-last messages as History (replayed through prefill) and sending the final user turn to trigger generation.

ChatHistory:  [System, User₁, Assistant₁, User₂]   (what SK hands the connector)
                 └──────── History ────────┘  └ Send ┘
                 (re-prefilled each call)     (generates)

This keeps SK's history and the model's KV cache in lockstep (SK owns the history and can edit it). The cost is an O(history) prefill per turn — fine for typical chats; for very long conversations, drive the native LiteRtConversation API directly. Calls are serialized (one live engine per process; conversations are not thread-safe), and the engine lifecycle is handled by the Extensions.AI registration the connector builds on. A ChatHistory system message is restored through that same History path, so this connector was never affected by the pre-v0.14.0 LiteRtConversationOptions.SystemMessage bug. And any LiteRtEngineOptions you pass to AddLiteRtChatCompletion, including the v0.14.0 additions (NumThreads, LoRA ranks), flows straight through.

Function calling

Set a FunctionChoiceBehavior on the execution settings and Semantic Kernel's functions are offered to the model; when the model calls one, it is auto-invoked and the result fed back so the model answers from it — the standard SK function-calling experience.

using System.ComponentModel;
using Microsoft.SemanticKernel;

// A plugin the model may call.
sealed class WeatherPlugin
{
    [KernelFunction, Description("Gets the current weather for a city.")]
    public string GetWeather([Description("The city.")] string city) => $"22°C and sunny in {city}";
}

kernel.Plugins.AddFromObject(new WeatherPlugin());

var settings = new LiteRtPromptExecutionSettings
{
    FunctionChoiceBehavior = FunctionChoiceBehavior.Auto(),
    EnableConstrainedDecoding = true,   // recommended for small models (see note); off by default
};

var answer = await kernel.InvokePromptAsync(
    "What's the weather in Paris? Use the weather tool.", new KernelArguments(settings));
Console.WriteLine(answer);              // "The weather in Paris is 22°C and sunny."

How it works. Semantic Kernel's AsChatCompletionService adapter passes the kernel's functions to the chat client as tools but does not run the auto-invoke loop itself, so the connector wraps the IChatClient with MEAI's function-invocation middleware (UseFunctionInvocation). That wrapping is built into AddLiteRtChatCompletion, so FunctionChoiceBehavior.Auto() just works; it is a no-op when a request carries no functions. Under the hood this is the same bridge the Microsoft.Extensions.AI integration exposes (the model's tool calls become FunctionCallContent).

Constrained decoding. LiteRtPromptExecutionSettings.EnableConstrainedDecoding = true makes the model emit valid, schema-shaped tool-call arguments — recommended for small on-device models. It is off by default and works on every platform (the linux-x64 restriction of earlier releases is gone with the official LiteRT-LM v0.16.0 prebuilts, which embed the constraint provider). Tools still work without it; arguments are just not grammar-constrained.

Function choice (FunctionChoiceBehavior). Auto() (the model decides), None() (no functions offered) and Required() (the model must call one) flow through to the chat client's ChatOptions.ToolMode. Required is best-effort on-device — the connector instructs the model to call a function and narrows the offered set, but cannot force a call the way a cloud API's server-enforced tool_choice: "required" does. See Tool choice for the per-mode behavior.

Structured output (response_format)

A raw JSON-Schema JsonElement under the standard response_format execution-settings key is enforced during sampling (native LiteRT-LM v0.15.0+): SK's adapter converts it into a Microsoft.Extensions.AI schema ResponseFormat, and the connector arms the LlGuidance constraint provider so the reply is guaranteed to be a conforming JSON document.

using JsonDocument schema = JsonDocument.Parse(
    """{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}""");

var settings = new LiteRtPromptExecutionSettings
{
    MaxTokens = 64,
    ExtensionData = new Dictionary<string, object> { ["response_format"] = schema.RootElement },
};
// result[0].Content parses as JSON conforming to the schema — enforced, not prompted.

Supported shapes (verified against a real model): the raw schema element shown above works; an OpenAI-style {"type":"json_schema","json_schema":{...}} envelope does not (SK forwards the whole envelope as the schema and the constraint compiler rejects it — the send fails); a plain "json_object" string is JSON mode, which is not enforced (prompt-driven only). The same rules and caveats as the MEAI connector apply — most importantly, not combinable with function calling on the same request (see the structured output section of the Extensions.AI guide).

Reasoning (thinking) and the output-token budget

With EnableThinking = true the model emits a reasoning trace before the answer, and that trace shares the MaxTokens output budget with the answer. So give thinking models headroom: with too small a budget the reasoning can consume it all and the answer comes back empty. For example, on gemma-4 the reasoning for a simple prompt runs ~200 tokens, so MaxTokens = 200 with thinking on leaves nothing for the answer, whereas MaxTokens = 512 works. This affects both blocking and streaming equally — it is a budget effect, not a connector bug.

A caveat specific to the Semantic Kernel path: SK's IChatCompletionService adapter does not surface the reasoning trace or a truncation signal — kernel.InvokePromptAsync / GetChatMessageContentsAsync return just the answer text (empty when the reasoning ate the budget). (Token usage, by contrast, does flow through — SK exposes it on the result's metadata; see Token usage.) The underlying Microsoft.Extensions.AI IChatClient does surface both: the reasoning as a TextReasoningContent (excluded from ChatResponse.Text), and a ChatResponse.FinishReason of Length when the answer was truncated. The same registration provides it, so resolve it from the kernel when you need rich reasoning / truncation handling:

using LiteRtLmSharp.Extensions.AI;   // for LiteRtChatOptions
using Microsoft.Extensions.AI;
using Microsoft.Extensions.DependencyInjection;

IChatClient chatClient = kernel.Services.GetRequiredService<IChatClient>();
ChatResponse response = await chatClient.GetResponseAsync(messages,
    new LiteRtChatOptions { MaxOutputTokens = 512, EnableThinking = true });

string? reasoning = response.Messages.SelectMany(m => m.Contents).OfType<TextReasoningContent>().FirstOrDefault()?.Text;
if (response.FinishReason == ChatFinishReason.Length && string.IsNullOrWhiteSpace(response.Text))
    Console.WriteLine("(no answer — the reasoning consumed the budget; raise MaxOutputTokens)");

Multimodal (image / audio)

On a multimodal model, add an ImageContent or AudioContent (with inline data) to a chat message — Semantic Kernel's adapter forwards it to the model as an attachment through the underlying IChatClient:

using Microsoft.SemanticKernel;
using Microsoft.SemanticKernel.ChatCompletion;

byte[] png = File.ReadAllBytes("photo.png");
var history = new ChatHistory();
history.Add(new ChatMessageContent(AuthorRole.User, new ChatMessageContentItemCollection
{
    new TextContent("What is in this image?"),
    new ImageContent(png, "image/png"),
}));

IChatCompletionService chat = kernel.GetRequiredService<IChatCompletionService>();
var reply = await chat.GetChatMessageContentsAsync(history);

The engine must have been loaded with the matching modality enabled (LiteRtEngineOptions.VisionBackend / AudioBackend) on a multimodal model. Only the final user message's media is sent (prior turns are restored as text). See docs/extensions-ai.md for the details.

Scope

  • Function calling is supported (see Function calling) via FunctionChoiceBehavior.
  • Multimodal (image / audio) is supported (see Multimodal).
  • Text generation (ITextGenerationService) is not provided — Semantic Kernel and the wider .NET AI stack are chat-centric; use chat completion.
  • Embeddings are not provided (the LiteRT-LM C API exposes none at v0.14.0).
  • Not AOT/trim-clean. Semantic Kernel itself is not; the core LiteRtLmSharp package stays AOT/trim-friendly, this companion does not carry that guarantee.

Sample

A runnable console sample is in samples/SemanticKernel: it builds a kernel with AddLiteRtChatCompletion and demonstrates a prompt function, a streaming prompt, a multi-turn streaming chat, and function calling (a [KernelFunction] plugin) — pass --interactive for a chat loop. See its README for how to run it. The broader Microsoft.Extensions.AI / Agent Framework story is in docs/extensions-ai.md.