Table of Contents

Class LiteRtEngineOptions

Namespace
LiteRtLmSharp
Assembly
LiteRtLmSharp.dll

Options for creating a LiteRtEngine.

public sealed record LiteRtEngineOptions : IEquatable<LiteRtEngineOptions>
Inheritance
LiteRtEngineOptions
Implements
Inherited Members

Properties

ActivationDataType

Activation tensor precision. null (default) uses the engine default (F16 for the text executor on GPU). Only the GPU backend honors this, and only as F32 vs F16: Float32 is higher precision at more memory and lower speed, Float16 is the faster default. On CPU it is a no-op, and Int16 / Int8 are accepted by the native API but not distinctly implemented by the shipped executors (folded to F16 on GPU). Maps to engine_settings_set_activation_data_type. See docs/engine-tuning.md.

public LiteRtActivationDataType? ActivationDataType { get; init; }

Property Value

LiteRtActivationDataType?

AudioBackend

Backend for the audio encoder, enabling audio input. null (default) leaves audio unconfigured — audio attachments will not work. Requires a model with audio support (e.g. the Gemma 4 E-series); a no-op on models without it.

public LiteRtBackend? AudioBackend { get; init; }

Property Value

LiteRtBackend?

Remarks

A model may constrain which backend its audio encoder accepts. The Gemma 4 audio sub-model requires CPU: passing Gpu makes litert_lm_engine_create fail with "Audio backend constraint mismatch. Model requires one of [cpu]" even when Backend is GPU — on any platform, not a win-x64 quirk. Use Cpu for such models; the vision encoder has no such constraint and runs on GPU. Maps to the audio_backend_str parameter of the C API engine_settings_create.

AudioLoraRank

LoRA rank for the audio executor. null (default) leaves the engine default. Only applies when an audio executor is configured (AudioBackend set). Maps to engine_settings_set_audio_lora_rank. See LoraRank for the not-yet-validated caveat.

public int? AudioLoraRank { get; init; }

Property Value

int?

Exceptions

ArgumentOutOfRangeException

The value is zero or negative.

AudioNumThreads

Number of CPU threads for the audio executor. null (default) leaves the engine default. Only applies when an audio executor is configured (AudioBackend set), and only reaches the audio backend's CPU compilation options — a no-op otherwise (verified in engine.cc set_audio_num_threads → audio_executor CPU options). Maps to engine_settings_set_audio_num_threads.

public int? AudioNumThreads { get; init; }

Property Value

int?

Exceptions

ArgumentOutOfRangeException

The value is zero or negative.

Backend

Backend to run the model on. Defaults to Cpu. Use Gpu, Npu, or Custom(string) for a backend exposed by your own native build.

public LiteRtBackend Backend { get; init; }

Property Value

LiteRtBackend

BenchmarkDecodeTokens

Synthetic decode token count for benchmarking — the decode half of BenchmarkPrefillTokens. 0 (default) = off. See BenchmarkPrefillTokens for the full behavior. Maps to engine_settings_set_num_decode_tokens.

public int BenchmarkDecodeTokens { get; init; }

Property Value

int

BenchmarkPrefillTokens

Synthetic prefill token count for benchmarking. 0 (default) = off (normal inference).

public int BenchmarkPrefillTokens { get; init; }

Property Value

int

Remarks

When > 0 the engine runs a synthetic benchmark instead of answering: the real prompt is truncated or padded to exactly this many tokens for prefill, and decoding runs exactly BenchmarkDecodeTokens tokens (ignoring the stop token). So GetBenchmarkInfo() reports prefill/decode throughput at FIXED token counts — independent of the prompt — which is useful for device throughput benchmarking and for measuring the effect of the tuning settings reproducibly. The reply text is not a real answer. Setting either this or BenchmarkDecodeTokens also turns benchmark mode on (the same switch as EnableBenchmark). Use a dedicated engine instance — do not reuse it for real chat. Verified observable through the Conversation API on win-x64 CPU (the default engine reads these during prefill/decode). Maps to engine_settings_set_num_prefill_tokens.

Cache

Where the engine keeps its compiled-artifact cache (GPU shaders / converted weights), which speeds up subsequent loads. Defaults to Default (written next to the model file). Use Disabled, InMemory, or Directory(string) for an explicit path.

public LiteRtCache Cache { get; init; }

Property Value

LiteRtCache

Remarks

Set this to Disabled to make EnableSpeculativeDecoding work on the desktop WebGPU GPU backend: with the default disk cache the MTP drafter's shared weight-cache file fails to open ("Access denied") on Windows and engine creation fails. This is an upstream issue (Google's own litert-lm CLI fails the same way with --cache disk and succeeds with --cache no); see docs/speculative-decoding.md.

EnableBenchmark

Enable benchmark instrumentation so GetBenchmarkInfo() returns prefill/decode tokens-per-second, time-to-first-token and init time. Fixed at engine creation; the overhead is timing bookkeeping only.

public bool EnableBenchmark { get; init; }

Property Value

bool

EnableSpeculativeDecoding

Enable speculative decoding — the model drafts several tokens ahead with a small Multi-Token-Prediction (MTP) drafter and the main model verifies them in one step, giving a large decode-throughput win (≈3× on supported models, per LiteRT-LM#2211).

public bool EnableSpeculativeDecoding { get; init; }

Property Value

bool

Remarks

Requires a .litertlm that ships an MTP drafter (e.g. the Gemma 4 E2B/E4B/12B builds). On a model without one the flag is a no-op (no speedup, no error). The setting is fixed at engine creation. Pair with EnableBenchmark to measure the gain (see GetBenchmarkInfo()).

Backend caveats (measured against LiteRT-LM v0.13.1 — see docs/speculative-decoding.md): the win comes from memory-bound accelerator decode. On desktop CPU it can REGRESS throughput (the drafter + verification overhead is not amortized). On the desktop WebGPU GPU backend it works, but only with the disk cache disabled — set Cache = Disabled, otherwise the drafter's shared weight-cache file fails to open ("Access denied") and engine creation fails (an upstream issue that reproduces in Google's own CLI).

EnableYnnpack

Lets the experimental YNNPACK delegate take the CPU operations it supports before XNNPACK. null (default) = engine default (off). CPU backend only. Requires native LiteRT-LM v0.16.0+.

public bool? EnableYnnpack { get; init; }

Property Value

bool?

Remarks

Maps to engine_settings_set_enable_ynnpack. Upstream ships the YNNPACK kernels in its linux-arm64 builds; the other official prebuilts accept the flag and run unchanged, so treat any speed-up as something to measure on the target device, not assume.

GpuDecodeStepsPerSync

Decode steps the GPU runs between host syncs. null (default) = engine default. Only honored by the Artisan GPU backend (the mobile GPU stack); the desktop WebGPU backend ignores it. Requires native LiteRT-LM v0.15.0+.

public int? GpuDecodeStepsPerSync { get; init; }

Property Value

int?

Remarks

Maps to engine_settings_set_gpu_decode_steps_per_sync. Larger values reduce sync overhead but coarsen streaming/cancellation granularity.

Exceptions

ArgumentOutOfRangeException

The value is zero or negative.

GpuWaitForWeightUploads

Whether engine load blocks until GPU weight uploads complete. null (default) = engine default. Only honored by the Artisan GPU backend; the desktop WebGPU backend ignores it. Requires native LiteRT-LM v0.15.0+.

public bool? GpuWaitForWeightUploads { get; init; }

Property Value

bool?

Remarks

Maps to engine_settings_set_gpu_wait_for_weight_uploads.

LoraRank

LoRA rank for the text executor. null (default) leaves the engine default (LoRA disabled). Requires a LoRA-enabled model and a matching adapter passed per conversation via LoraPath. Maps to engine_settings_set_lora_rank.

public int? LoraRank { get; init; }

Property Value

int?

Remarks

The full LoRA path (rank here + adapter file on the conversation) has not yet been validated end-to-end in this binding — no LoRA adapter artifact was available at the time of writing — so treat it as wired-through but unverified against a real adapter.

Exceptions

ArgumentOutOfRangeException

The value is zero or negative.

MaxNumImages

Maximum number of images the engine accepts per turn. 0 = engine default. On the standard binaries this setting has no effect — see the remarks for the case where it applies. To bound how much of the context window images consume, use VisualTokenBudget instead.

public int MaxNumImages { get; init; }

Property Value

int

Remarks

The native library compiles several engine implementations into one binary and picks one per backend. This value always reaches the engine settings, but only the legacy TFLite implementation reads it; the standard binaries select the modern CompiledModel engines for CPU/GPU, which ignore it. It matters only when a custom native build routes your backend (e.g. one targeted via Custom(string)) through the legacy engine. Kept for parity with the reference Kotlin binding's maxNumImages. Maps to engine_settings_set_max_num_images.

MaxNumTokens

The total context window in tokens: prompt + generated replies, accumulated across every turn of a conversation. 0 = engine default.

public int MaxNumTokens { get; init; }

Property Value

int

Remarks

This is the hard ceiling for a conversation's KV cache. Size it for the longest exchange you expect — the system prompt, any restored History, and each user turn and reply all count against it and accumulate. As the running total (TokenCount) nears the limit, generation degrades, so manage the conversation before then: trim history, cap replies with MaxOutputTokens, or start a fresh conversation. A larger window costs more memory and a slower prefill, so prefer the smallest that fits your use case. For multimodal input leave room for the media on top of the text — an image expands to roughly 256 vision tokens — for which 4096 is a comfortable starting point.

Setting this also arms the binding's KV overflow guard. The native runtime does not police the limit: a send that grows past it writes beyond the allocated cache and corrupts native memory (typically a deferred 0xC0000005 process crash on a later call). With an explicit value here, the send methods instead clamp each reply to the remaining context and throw LiteRtContextOverflowException when the conversation is full. The clamp cannot cover sends whose prefill cost is unmeasurable managed-side — media attachments and SendRaw's per-send extraContext get only the conversation-full check, so keep extra headroom under the limit on those paths (an image is roughly 256 tokens). With 0 the effective limit is internal to the engine (the C API exposes no getter), so the guard stays off.

Keep the value at or below the model's built-in context length. The native loader sizes the KV cache from this hint only when it is smaller than the model's own maximum; a value at or above it silently falls back to the model's default sizing, leaving the real cache SMALLER than this property claims — and the overflow guard, calibrated against this value, would then stop policing before the actual edge. Values below the model's minimum prefill work group (~128) are unusable outright: every send is rejected.

Keep the value at or above the model's largest prefill signature (1024 for the published gemma conversions). The native loader accepts a smaller limit, but a send whose prefill spans more than the smallest work group then fails inside the native graph (an internal DYNAMIC_UPDATE_SLICE error) instead of cleanly. The C API exposes no way to query the signatures, so the binding cannot validate this up front; when a send fails on an engine loaded with a limit below 1024, the exception names this as the likely cause.

ModelPath

Path to the .litertlm (or .task) model file. Must be set — Load(LiteRtEngineOptions) throws ArgumentException when it is empty. (Deliberately not required: future native versions add alternate model sources — e.g. loading from a file descriptor — which will land here as sibling properties.)

public string ModelPath { get; init; }

Property Value

string

NumThreads

Number of CPU threads for the text executor. null (default) leaves the engine default. CPU-backend only: the native setter reads the executor's CpuConfig and is a no-op when the text backend is not CPU (verified in engine.cc set_num_threads). Maps to engine_settings_set_num_threads.

public int? NumThreads { get; init; }

Property Value

int?

Exceptions

ArgumentOutOfRangeException

The value is zero or negative.

ParallelFileSectionLoading

Whether the engine loads the .litertlm file's sections in parallel during startup. null (default) leaves the engine default (on). When on, the tokenizer section is parsed on a background thread while the model is built, shortening cold-start init; set false to load it serially on the calling thread (single-threaded environments, or to avoid the brief concurrent init peak, at the cost of a slower start). Maps to engine_settings_set_parallel_file_section_loading. See docs/engine-tuning.md.

public bool? ParallelFileSectionLoading { get; init; }

Property Value

bool?

PrefillChunkSize

Maximum prompt tokens prefilled per step. 0 (default) = no chunking (the whole prompt is prefilled at once). CPU + dynamic models only — ignored on GPU and on static models. A smaller chunk lowers peak memory during prefill and allows more timely cancellation of a long prompt, at the cost of more prefill iterations (potentially slower). Maps to engine_settings_set_prefill_chunk_size. See docs/engine-tuning.md.

public int PrefillChunkSize { get; init; }

Property Value

int

SupportedAudioLoraRanks

The set of LoRA ranks the audio executor should support. null (default) leaves the engine default. Requires an audio executor (AudioBackend set); see SupportedLoraRanks for the GPU-backend caveat. Every element must be positive and the list non-empty. Maps to engine_settings_set_supported_audio_lora_ranks.

public IReadOnlyList<int>? SupportedAudioLoraRanks { get; init; }

Property Value

IReadOnlyList<int>

Exceptions

ArgumentException

The list is empty.

ArgumentOutOfRangeException

Any element is zero or negative.

SupportedLoraRanks

The set of LoRA ranks the text executor should support (for switching adapters of different ranks). null (default) leaves the engine default. Only honored on the GPU (Artisan) backend: on other backends the native layer accepts and silently ignores it (verified in llm_executor_settings.h SetSupportedLoraRanks). Every element must be positive and the list non-empty. Maps to engine_settings_set_supported_lora_ranks.

public IReadOnlyList<int>? SupportedLoraRanks { get; init; }

Property Value

IReadOnlyList<int>

Exceptions

ArgumentException

The list is empty.

ArgumentOutOfRangeException

Any element is zero or negative.

UseRingbuffersLocalAttention

Store the KV cache of local-attention layers in ringbuffers sized to what those layers actually need, instead of allocating the full context length — lower memory for long contexts, at the cost of instant rewinding. null (default) = off. Backend-agnostic in interface but currently only the Artisan GPU backend implements it; unsupported models/backends log a native warning and ignore it. Requires native LiteRT-LM v0.15.0+.

public bool? UseRingbuffersLocalAttention { get; init; }

Property Value

bool?

Remarks

Maps to engine_settings_set_use_ringbuffers_local_attention (the knob the JS API exposes as use_autosized_ringbuffers).

VisionBackend

Backend for the vision encoder, enabling image input. null (default) leaves vision unconfigured — image attachments will not work. Requires a multimodal model (e.g. the Gemma 4 E-series); on a text-only model setting this has no effect.

public LiteRtBackend? VisionBackend { get; init; }

Property Value

LiteRtBackend?

Remarks

Maps to the vision_backend_str parameter of the C API engine_settings_create. May differ from Backend (e.g. main on Gpu, vision on Cpu). Pair with image attachments via Image(ReadOnlySpan<byte>) and tune the image prefill budget with VisualTokenBudget.