Table of Contents

Namespace LiteRtLmSharp

Classes

LiteRtAttachment

One image or audio attachment to add to a user message on a multimodal model. Attach it to a turn with Send(string, IReadOnlyList<LiteRtAttachment>?, LiteRtSendOptions?) (or the streaming overload). The engine must have been loaded with the matching modality enabled (VisionBackend for images, AudioBackend for audio) and the model must support it (e.g. the Gemma 4 E-series handle both).

Two carriers, both mapping to the LiteRT-LM content-part wire format ({"type":"image"|"audio", …}): inline LiteRtLmSharp.LiteRtAttachment.Bytes (sent as a base64 "blob"), or a LiteRtLmSharp.LiteRtAttachment.Path on disk (sent as "path" and memory-mapped natively, avoiding a base64 round-trip — desktop only; the path must be readable by the native process).

LiteRtBenchmarkInfo

A snapshot of the engine's benchmark counters for one conversation, returned by GetBenchmarkInfo(). Only populated when the engine was created with EnableBenchmark = true.

Prefill and decode metrics accumulate one entry per turn; the Last* properties report the most recent turn (the headline number for "how fast was that reply"). Decode throughput is the metric that EnableSpeculativeDecoding accelerates.

LiteRtConstraint

A per-send output constraint enforced during sampling by the conversation's ConstraintProvider: tokens that cannot lead to a valid completion of the pattern are masked out, so the model can only emit conforming output. Set on Constraint; requires the conversation to have been created with LlGuidance. Requires native LiteRT-LM v0.15.0+.

LiteRtContextOverflowException

Raised instead of letting a send overflow the conversation's KV cache. The KV cache is sized by MaxNumTokens and the native runtime does not police it: a send that grows past the limit writes beyond the allocated cache, corrupting native memory — the typical symptom is a deferred process crash (0xC0000005 / 0xC0000374) on a LATER call, with nothing pointing back at the overflowing send. The binding throws this instead, before native state is damaged.

LiteRtConversation

A stateful conversation over a LiteRtEngine. Handles chat templating internally (mirrors the Gemini Chat APIs). Not thread-safe: serialize calls per instance — and serialize sends per engine, not just per conversation: multiple live conversations may be interleaved freely, but concurrent generation on different conversations of one engine is not supported by the native runtime (non-deterministic output and GPU-backend errors, upstream LiteRT-LM#2807; its thread-safety contract is undocumented).

LiteRtConversationOptions

Per-conversation options: system prompt, sampler, output limit, and tools (function calling). Requires native binaries version-matched to the bindings (the official builds from this repo's releases are); see docs/native-abi.md for the ABI history.

LiteRtEngine

A LiteRT-LM inference engine. Heavyweight: holds the model weights. Create one per model and spawn lightweight LiteRtConversation objects from it.

LiteRtEngineOptions

Options for creating a LiteRtEngine.

LiteRtException

Raised when a LiteRT-LM native call fails.

LiteRtMessage

One message in a conversation history, used to restore a conversation: pass a list of these as History and the engine re-prefills them when the conversation is created. Build them with the factory methods (User(string), Model(string, IReadOnlyList<LiteRtToolCall>?), Tool(IEnumerable<LiteRtToolResult>), System(string)); capture an assistant turn from a live reply with ToMessage(bool).

The LiteRT-LM C API exposes no way to read a conversation's history back, so the round-trip is caller-owned: record each turn as it happens, persist with Serialize(IEnumerable<LiteRtMessage>), and reload with Deserialize(string) (or feed the serialized JSON straight to HistoryJson). Restoring replays the history through prefill; it is not a zero-cost KV-cache snapshot.

LiteRtNoRepeatNgramOptions

Bans the current send's decode from repeating any n-gram of NgramSize consecutive tokens it already produced: a candidate token that would complete an already-seen n-gram has its logit forced to -inf. Set on NoRepeatNgram. Requires native LiteRT-LM v0.15.0+.

LiteRtRepetitionPenaltyOptions

Repetition penalties for one send's decode, in the two industry conventions the native runtime implements: a multiplicative HuggingFace-style RepetitionPenalty and subtractive OpenAI-style PresencePenalty/FrequencyPenalty. All penalties look at the generated output window of the current send only (not the prompt, not prior turns). Set on RepetitionPenalties. Requires native LiteRT-LM v0.15.0+.

LiteRtResponse

A model response: plain Text or one or more ToolCalls, plus — on reasoning models — a separate Thinking trace (see Channels). RawJson always holds the unparsed response (escape hatch).

LiteRtSamplerParams

Sampler parameters for a conversation. Construct one and set it on Sampler. A constructed instance sends all fields to the engine — replacing any sampler parameters the model file itself carries — so set the ones you care about and leave the rest at their defaults. The defaults below (TopP, 40, 0.95, 1.0, seed 0) are the same values Google's official bindings fill in for unset fields, so a partial configuration behaves identically across the LiteRT-LM ecosystem.

LiteRtSendOptions

Per-send options for Send(string, IReadOnlyList<LiteRtAttachment>?, LiteRtSendOptions?) and its async/streaming variants. Everything here applies to one send, overriding the conversation-level setting where one exists; pass null (the default) to use the conversation's configuration unchanged.

LiteRtTokenUnion

A single model start (BOS) or stop (EOS) token, as reported by GetStartToken() / GetStopTokens(). A model expresses each such token either as a literal Text string or as a sequence of token Ids; Kind says which. Use it to recognize when generated output reaches a stop token, or to compare against Tokenize(string) output.

LiteRtTool

A tool the model may call. ParametersJson is a JSON-Schema object string (the parameters of a Gemini/OpenAI function declaration), e.g. {"type":"object","properties":{"location":{"type":"string"}},"required":["location"]}.

LiteRtToolCall

A tool/function call emitted by the model. ArgumentsJson is a raw JSON object.

LiteRtToolResult

The result of executing a tool, to be sent back to the model. ResultJson is a raw JSON value (object/array/scalar).

Structs

LiteRtBackend

A compute backend for the engine or one of its encoders. Use a well-known value (Cpu, Gpu, Npu) or Custom(string) for a backend exposed by your own native build. The underlying native identifier is Value; backends compare by it.

LiteRtCache

Where the engine keeps its compiled-artifact cache (GPU shaders / converted weights), which speeds up subsequent loads. Choose Default (next to the model file), Disabled, InMemory, or Directory(string) for an explicit path. Set it on Cache.

LiteRtStreamChunk

One streamed piece of a reply from SendStreamingAsync(string, CancellationToken). Kind says what it is: an Answer or Thinking text delta (concatenate same-kind chunks in order to rebuild each), a ToolCall carrying the model's tool calls, or an opt-in ToolCallDelta raw progress fragment. Text is the delta for the text kinds (empty for complete tool calls); ToolCalls is populated only for the tool-call kind.

Enums

LiteRtActivationDataType

Activation tensor precision for ActivationDataType, mirroring upstream's ActivationDataType enum. Only Float32 and Float16 are distinctly honored, and only on the GPU backend (see ActivationDataType).

LiteRtAttachmentKind

The media kind carried by a LiteRtAttachment.

LiteRtConstraintProvider

Provider used to compile and enforce custom LiteRtConstraints (the "custom constrained decoding" path). Distinct from EnableConstrainedDecoding, which is the tool-calling constrained-decoding path — the native runtime supports only one of the two per conversation.

LiteRtConstraintType

The shape of a LiteRtConstraint pattern string.

LiteRtMessageRole

The author of a LiteRtMessage in a conversation history.

LiteRtSamplerType

Token sampling strategy. Mirrors LiteRT-LM's LiteRtLmSamplerType.

LiteRtStreamChunkKind

What a LiteRtStreamChunk carries.

LiteRtTokenKind

Whether a LiteRtTokenUnion carries a literal string or a sequence of token ids.