LiteRtLmSharp
.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) for any .NET app, including MAUI. P/Invoke over LiteRT-LM's C API, with native binaries distributed per-RID as NuGet packages.
Chat (blocking, awaitable + cancellable, streaming), function calling, multimodal (image/audio), conversation restore/clone, reasoning mode, tokenizer, speculative decoding and benchmarking — all running locally, no server.
The same MAUI sample app, on an Android phone and on Windows (GPU) — both fully on-device.
Install
<PackageReference Include="LiteRtLmSharp" Version="1.1.1" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.1.1" />
<!-- or LiteRtLmSharp.runtime.linux-x64 / android-arm64 / osx-arm64, per target -->
Install the managed package plus the runtime package for your platform, always with the same version
number. Optional integrations:
LiteRtLmSharp.Extensions.AI (IChatClient — Microsoft Agent Framework, MEAI) and
LiteRtLmSharp.SemanticKernel (IChatCompletionService).
First tokens
using LiteRtLmSharp;
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", // from huggingface.co/litert-community
Backend = LiteRtBackend.Cpu, // or .Gpu (WebGPU -> D3D12/Vulkan/Metal)
MaxNumTokens = 4096, // total context window
});
using var chat = engine.CreateConversation();
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
Console.Write(chunk.Text);
The repository README walks through every feature with runnable snippets; the guides here go deeper per topic.
Where to go next
- Chat & generation — function calling, reasoning mode, multimodal (image/audio), token counting.
- Microsoft.Extensions.AI integration — plug the model into the .NET AI
ecosystem as an
IChatClient(Agent Framework, Semantic Kernel, MEAI middleware). - Semantic Kernel connector —
IChatCompletionServiceover the same bridge. - Conversation state — persist/restore chats, clone live conversations.
- Engine tuning — precision, prefill chunking, thread counts, benchmarking.
- Speculative decoding — the MTP drafter: when it helps and what it needs.
- Android — device setup, backends, and MAUI notes.
- API Reference — the full public surface, generated from the XML docs.
Samples
samples/Console— chat loop with--tools,--spec, and--thinkingdemos.samples/Maui— full Android/Windows chat app: model download, streaming, multimodal attachments, function calling.samples/SemanticKernel— kernel registration, streaming, and a[KernelFunction]plugin.
Project status
The roadmap is the status source of truth; release history is in the changelog.