Class LiteRtEngineOptions
- Namespace
- LiteRtLmSharp
- Assembly
- LiteRtLmSharp.dll
Options for creating a LiteRtEngine.
public sealed record LiteRtEngineOptions : IEquatable<LiteRtEngineOptions>
- Inheritance
-
LiteRtEngineOptions
- Implements
- Inherited Members
Properties
ActivationDataType
Activation tensor precision. null (default) uses the engine default (F16 for the text
executor on GPU). Only the GPU backend honors this, and only as F32 vs F16:
Float32 is higher precision at more memory and lower speed,
Float16 is the faster default. On CPU it is a no-op,
and Int16 / Int8 are
accepted by the native API but not distinctly implemented by the shipped executors (folded to F16
on GPU). Maps to engine_settings_set_activation_data_type. See docs/engine-tuning.md.
public LiteRtActivationDataType? ActivationDataType { get; init; }
Property Value
AudioBackend
Backend for the audio encoder, enabling audio input. null (default) leaves audio
unconfigured — audio attachments will not work. Requires a model with audio support (e.g. the
Gemma 4 E-series); a no-op on models without it.
public LiteRtBackend? AudioBackend { get; init; }
Property Value
Remarks
A model may constrain which backend its audio encoder accepts. The Gemma 4 audio sub-model
requires CPU: passing Gpu makes litert_lm_engine_create
fail with "Audio backend constraint mismatch. Model requires one of [cpu]" even when
Backend is GPU — on any platform, not a win-x64 quirk. Use
Cpu for such models; the vision encoder has no such constraint and
runs on GPU. Maps to the audio_backend_str parameter of the C API
engine_settings_create.
AudioLoraRank
LoRA rank for the audio executor. null (default) leaves the engine default. Only
applies when an audio executor is configured (AudioBackend set). Maps to
engine_settings_set_audio_lora_rank. See LoraRank for the not-yet-validated
caveat.
public int? AudioLoraRank { get; init; }
Property Value
- int?
Exceptions
- ArgumentOutOfRangeException
The value is zero or negative.
AudioNumThreads
Number of CPU threads for the audio executor. null (default) leaves the engine
default. Only applies when an audio executor is configured (AudioBackend set), and
only reaches the audio backend's CPU compilation options — a no-op otherwise (verified in
engine.cc set_audio_num_threads → audio_executor CPU options). Maps to
engine_settings_set_audio_num_threads.
public int? AudioNumThreads { get; init; }
Property Value
- int?
Exceptions
- ArgumentOutOfRangeException
The value is zero or negative.
Backend
Backend to run the model on. Defaults to Cpu. Use Gpu, Npu, or Custom(string) for a backend exposed by your own native build.
public LiteRtBackend Backend { get; init; }
Property Value
BenchmarkDecodeTokens
Synthetic decode token count for benchmarking — the decode half of
BenchmarkPrefillTokens. 0 (default) = off. See BenchmarkPrefillTokens
for the full behavior. Maps to engine_settings_set_num_decode_tokens.
public int BenchmarkDecodeTokens { get; init; }
Property Value
BenchmarkPrefillTokens
Synthetic prefill token count for benchmarking. 0 (default) = off (normal inference).
public int BenchmarkPrefillTokens { get; init; }
Property Value
Remarks
When > 0 the engine runs a synthetic benchmark instead of answering: the real prompt is
truncated or padded to exactly this many tokens for prefill, and decoding runs exactly
BenchmarkDecodeTokens tokens (ignoring the stop token). So
GetBenchmarkInfo() reports prefill/decode throughput at FIXED token
counts — independent of the prompt — which is useful for device throughput benchmarking and for
measuring the effect of the tuning settings reproducibly. The reply text is not a real
answer. Setting either this or BenchmarkDecodeTokens also turns benchmark mode on
(the same switch as EnableBenchmark). Use a dedicated engine instance — do not reuse
it for real chat. Verified observable through the Conversation API on win-x64 CPU (the default
engine reads these during prefill/decode). Maps to engine_settings_set_num_prefill_tokens.
Cache
Where the engine keeps its compiled-artifact cache (GPU shaders / converted weights), which speeds up subsequent loads. Defaults to Default (written next to the model file). Use Disabled, InMemory, or Directory(string) for an explicit path.
public LiteRtCache Cache { get; init; }
Property Value
Remarks
Set this to Disabled to make EnableSpeculativeDecoding
work on the desktop WebGPU GPU backend: with the default disk cache the MTP drafter's
shared weight-cache file fails to open ("Access denied") on Windows and engine creation fails.
This is an upstream issue (Google's own litert-lm CLI fails the same way with
--cache disk and succeeds with --cache no); see docs/speculative-decoding.md.
EnableBenchmark
Enable benchmark instrumentation so GetBenchmarkInfo() returns prefill/decode tokens-per-second, time-to-first-token and init time. Fixed at engine creation; the overhead is timing bookkeeping only.
public bool EnableBenchmark { get; init; }
Property Value
EnableSpeculativeDecoding
Enable speculative decoding — the model drafts several tokens ahead with a small Multi-Token-Prediction (MTP) drafter and the main model verifies them in one step, giving a large decode-throughput win (≈3× on supported models, per LiteRT-LM#2211).
public bool EnableSpeculativeDecoding { get; init; }
Property Value
Remarks
Requires a .litertlm that ships an MTP drafter (e.g. the Gemma 4 E2B/E4B/12B
builds). On a model without one the flag is a no-op (no speedup, no error). The setting
is fixed at engine creation. Pair with EnableBenchmark to measure the gain
(see GetBenchmarkInfo()).
Backend caveats (measured against LiteRT-LM v0.13.1 — see
docs/speculative-decoding.md): the win comes from memory-bound accelerator decode.
On desktop CPU it can REGRESS throughput (the drafter + verification overhead is not
amortized). On the desktop WebGPU GPU backend it works, but only with the disk cache
disabled — set Cache = Disabled, otherwise the drafter's
shared weight-cache file fails to open ("Access denied") and engine creation fails (an
upstream issue that reproduces in Google's own CLI).
EnableYnnpack
Lets the experimental YNNPACK delegate take the CPU operations it supports before XNNPACK.
null (default) = engine default (off). CPU backend only. Requires native LiteRT-LM v0.16.0+.
public bool? EnableYnnpack { get; init; }
Property Value
- bool?
Remarks
Maps to engine_settings_set_enable_ynnpack. Upstream ships the YNNPACK kernels in its
linux-arm64 builds; the other official prebuilts accept the flag and run unchanged, so treat any
speed-up as something to measure on the target device, not assume.
GpuDecodeStepsPerSync
Decode steps the GPU runs between host syncs. null (default) = engine default. Only
honored by the Artisan GPU backend (the mobile GPU stack); the desktop WebGPU backend ignores
it. Requires native LiteRT-LM v0.15.0+.
public int? GpuDecodeStepsPerSync { get; init; }
Property Value
- int?
Remarks
Maps to engine_settings_set_gpu_decode_steps_per_sync. Larger values reduce
sync overhead but coarsen streaming/cancellation granularity.
Exceptions
- ArgumentOutOfRangeException
The value is zero or negative.
GpuWaitForWeightUploads
Whether engine load blocks until GPU weight uploads complete. null (default) = engine
default. Only honored by the Artisan GPU backend; the desktop WebGPU backend ignores it.
Requires native LiteRT-LM v0.15.0+.
public bool? GpuWaitForWeightUploads { get; init; }
Property Value
- bool?
Remarks
Maps to engine_settings_set_gpu_wait_for_weight_uploads.
LoraRank
LoRA rank for the text executor. null (default) leaves the engine default (LoRA
disabled). Requires a LoRA-enabled model and a matching adapter passed per conversation via
LoraPath. Maps to engine_settings_set_lora_rank.
public int? LoraRank { get; init; }
Property Value
- int?
Remarks
The full LoRA path (rank here + adapter file on the conversation) has not yet been validated end-to-end in this binding — no LoRA adapter artifact was available at the time of writing — so treat it as wired-through but unverified against a real adapter.
Exceptions
- ArgumentOutOfRangeException
The value is zero or negative.
MaxNumImages
Maximum number of images the engine accepts per turn. 0 = engine default. On the standard binaries this setting has no effect — see the remarks for the case where it applies. To bound how much of the context window images consume, use VisualTokenBudget instead.
public int MaxNumImages { get; init; }
Property Value
Remarks
The native library compiles several engine implementations into one binary and picks one per
backend. This value always reaches the engine settings, but only the legacy TFLite
implementation reads it; the standard binaries select the modern CompiledModel engines for
CPU/GPU, which ignore it. It matters only when a custom native build routes your backend
(e.g. one targeted via Custom(string)) through the legacy engine. Kept for
parity with the reference Kotlin binding's maxNumImages. Maps to
engine_settings_set_max_num_images.
MaxNumTokens
The total context window in tokens: prompt + generated replies, accumulated across every turn of a conversation. 0 = engine default.
public int MaxNumTokens { get; init; }
Property Value
Remarks
This is the hard ceiling for a conversation's KV cache. Size it for the longest exchange you expect — the system prompt, any restored History, and each user turn and reply all count against it and accumulate. As the running total (TokenCount) nears the limit, generation degrades, so manage the conversation before then: trim history, cap replies with MaxOutputTokens, or start a fresh conversation. A larger window costs more memory and a slower prefill, so prefer the smallest that fits your use case. For multimodal input leave room for the media on top of the text — an image expands to roughly 256 vision tokens — for which 4096 is a comfortable starting point.
Setting this also arms the binding's KV overflow guard. The native runtime does not
police the limit: a send that grows past it writes beyond the allocated cache and corrupts native
memory (typically a deferred 0xC0000005 process crash on a later call). With an explicit
value here, the send methods instead clamp each reply to the remaining context and throw
LiteRtContextOverflowException when the conversation is full. The clamp cannot
cover sends whose prefill cost is unmeasurable managed-side — media attachments and
SendRaw's per-send extraContext get only the conversation-full check, so keep extra
headroom under the limit on those paths (an image is roughly 256 tokens). With 0 the effective
limit is internal to the engine (the C API exposes no getter), so the guard stays off.
Keep the value at or below the model's built-in context length. The native loader sizes the KV cache from this hint only when it is smaller than the model's own maximum; a value at or above it silently falls back to the model's default sizing, leaving the real cache SMALLER than this property claims — and the overflow guard, calibrated against this value, would then stop policing before the actual edge. Values below the model's minimum prefill work group (~128) are unusable outright: every send is rejected.
Keep the value at or above the model's largest prefill signature (1024 for the
published gemma conversions). The native loader accepts a smaller limit, but a send whose prefill
spans more than the smallest work group then fails inside the native graph (an internal
DYNAMIC_UPDATE_SLICE error) instead of cleanly. The C API exposes no way to query the
signatures, so the binding cannot validate this up front; when a send fails on an engine loaded
with a limit below 1024, the exception names this as the likely cause.
ModelPath
Path to the .litertlm (or .task) model file. Must be set —
Load(LiteRtEngineOptions) throws ArgumentException when it is empty.
(Deliberately not required: future native versions add alternate model sources — e.g.
loading from a file descriptor — which will land here as sibling properties.)
public string ModelPath { get; init; }
Property Value
NumThreads
Number of CPU threads for the text executor. null (default) leaves the engine
default. CPU-backend only: the native setter reads the executor's CpuConfig and is a
no-op when the text backend is not CPU (verified in engine.cc set_num_threads). Maps to
engine_settings_set_num_threads.
public int? NumThreads { get; init; }
Property Value
- int?
Exceptions
- ArgumentOutOfRangeException
The value is zero or negative.
ParallelFileSectionLoading
Whether the engine loads the .litertlm file's sections in parallel during startup.
null (default) leaves the engine default (on). When on, the tokenizer section is parsed on
a background thread while the model is built, shortening cold-start init; set false to load
it serially on the calling thread (single-threaded environments, or to avoid the brief concurrent
init peak, at the cost of a slower start). Maps to
engine_settings_set_parallel_file_section_loading. See docs/engine-tuning.md.
public bool? ParallelFileSectionLoading { get; init; }
Property Value
- bool?
PrefillChunkSize
Maximum prompt tokens prefilled per step. 0 (default) = no chunking (the whole prompt is
prefilled at once). CPU + dynamic models only — ignored on GPU and on static models. A
smaller chunk lowers peak memory during prefill and allows more timely cancellation of a long
prompt, at the cost of more prefill iterations (potentially slower). Maps to
engine_settings_set_prefill_chunk_size. See docs/engine-tuning.md.
public int PrefillChunkSize { get; init; }
Property Value
SupportedAudioLoraRanks
The set of LoRA ranks the audio executor should support. null (default) leaves the
engine default. Requires an audio executor (AudioBackend set); see
SupportedLoraRanks for the GPU-backend caveat. Every element must be positive and the
list non-empty. Maps to engine_settings_set_supported_audio_lora_ranks.
public IReadOnlyList<int>? SupportedAudioLoraRanks { get; init; }
Property Value
Exceptions
- ArgumentException
The list is empty.
- ArgumentOutOfRangeException
Any element is zero or negative.
SupportedLoraRanks
The set of LoRA ranks the text executor should support (for switching adapters of different
ranks). null (default) leaves the engine default. Only honored on the GPU (Artisan)
backend: on other backends the native layer accepts and silently ignores it (verified in
llm_executor_settings.h SetSupportedLoraRanks). Every element must be positive and the list
non-empty. Maps to engine_settings_set_supported_lora_ranks.
public IReadOnlyList<int>? SupportedLoraRanks { get; init; }
Property Value
Exceptions
- ArgumentException
The list is empty.
- ArgumentOutOfRangeException
Any element is zero or negative.
UseRingbuffersLocalAttention
Store the KV cache of local-attention layers in ringbuffers sized to what those layers
actually need, instead of allocating the full context length — lower memory for long
contexts, at the cost of instant rewinding. null (default) = off. Backend-agnostic in
interface but currently only the Artisan GPU backend implements it; unsupported
models/backends log a native warning and ignore it. Requires native LiteRT-LM v0.15.0+.
public bool? UseRingbuffersLocalAttention { get; init; }
Property Value
- bool?
Remarks
Maps to engine_settings_set_use_ringbuffers_local_attention (the knob the JS
API exposes as use_autosized_ringbuffers).
VisionBackend
Backend for the vision encoder, enabling image input. null (default) leaves vision
unconfigured — image attachments will not work. Requires a multimodal model (e.g. the Gemma 4
E-series); on a text-only model setting this has no effect.
public LiteRtBackend? VisionBackend { get; init; }
Property Value
Remarks
Maps to the vision_backend_str parameter of the C API engine_settings_create.
May differ from Backend (e.g. main on Gpu, vision on
Cpu). Pair with image attachments via
Image(ReadOnlySpan<byte>) and tune the image prefill budget
with VisualTokenBudget.