Table of Contents

Class LiteRtConversationOptions

Namespace
LiteRtLmSharp
Assembly
LiteRtLmSharp.dll

Per-conversation options: system prompt, sampler, output limit, and tools (function calling). Requires native binaries version-matched to the bindings (the official builds from this repo's releases are); see docs/native-abi.md for the ABI history.

public sealed record LiteRtConversationOptions : IEquatable<LiteRtConversationOptions>
Inheritance
LiteRtConversationOptions
Implements
Inherited Members

Properties

AudioLoraPath

Path to an audio LoRA adapter (weights file) to apply to this conversation. null (default) = no adapter. Requires an audio-capable, LoRA-enabled model and a matching AudioLoraRank. Opened at conversation creation (see LoraPath for the fail-fast and not-yet-validated notes). Maps to the C API session_config_set_audio_lora_path.

public string? AudioLoraPath { get; init; }

Property Value

string

ConstraintProvider

Enables custom constrained decoding for this conversation: with a provider set, each send may carry a Constraint (regex / JSON Schema) that the provider enforces during sampling. null (default) = no custom constraints.

public LiteRtConstraintProvider? ConstraintProvider { get; init; }

Property Value

LiteRtConstraintProvider?

Remarks

Mutually exclusive with EnableConstrainedDecoding (the tool-calling constrained-decoding path) — the native runtime supports one per conversation, and CreateConversation(LiteRtConversationOptions?) throws ArgumentException when both are set. Maps to the C API conversation_config_set_constraint_provider (native v0.15.0+). The LlGuidance provider, like the tool-calling one, is embedded in the native library.

EnableConstrainedDecoding

Force the model to emit valid (schema-constrained) output. Strongly recommended when Tools are set so tool-call arguments parse reliably.

public bool EnableConstrainedDecoding { get; init; }

Property Value

bool

Remarks

Supported on every platform with the official LiteRT-LM v0.16.0 prebuilts, whose constraint provider is embedded in the native library. Releases before 1.2.0 threw PlatformNotSupportedException on linux-x64, where the separately shipped provider crashed the process (google-ai-edge/LiteRT-LM#2149); that guard is gone. Tools also work with this left false; arguments are then not grammar-constrained.

EnableThinking

Toggles the model's reasoning ("thinking") mode by setting enable_thinking in the conversation's ExtraContext. null (default) leaves it unset so the model uses its own default; true/false force reasoning on/off.

public bool? EnableThinking { get; init; }

Property Value

bool?

Remarks

This is the canonical use of extra context for the Gemma reasoning builds: the chat template branches on {% if enable_thinking %}. On a model whose template does not reference the key it is a harmless no-op. Pair with FilterThinkingFromKvCache to keep the (often long) reasoning out of the KV cache on later turns. When both this and ExtraContext set enable_thinking, this flag wins.

ExtraContext

Extra context merged into the conversation preface and passed to the prompt-template renderer — a raw JSON object string, e.g. {"user_name":"Alice"}. Template variables are referenced Jinja-style ({{ user_name }}). Null/empty = none.

public string? ExtraContext { get; init; }

Property Value

string

Remarks

For the common enable_thinking toggle prefer the typed EnableThinking; this is the general escape hatch for arbitrary template variables. Must be a JSON object — CreateConversation(LiteRtConversationOptions?) throws ArgumentException otherwise (validated when the conversation is created, not when this property is set). Maps to the C API conversation_config_set_extra_context.

FilterThinkingFromKvCache

Drop channel content — in practice the thinking channel — from the KV cache, so a long reasoning block does not eat into the context window on subsequent turns. Default false. Only meaningful alongside EnableThinking.

public bool FilterThinkingFromKvCache { get; init; }

Property Value

bool

Remarks

Maps to the C API conversation_config_set_filter_channel_content_from_kv_cache.

History

Conversation history to restore (or a few-shot preface) — the prior turns are re-prefilled into the KV cache when the conversation is created, so the model continues as if they had just happened. Build the list with LiteRtMessage factories and capture assistant turns with ToMessage(bool); persist and reload via Serialize(IEnumerable<LiteRtMessage>) / Deserialize(string).

public IReadOnlyList<LiteRtMessage>? History { get; init; }

Property Value

IReadOnlyList<LiteRtMessage>

Remarks

The C API has no history getter, so the round-trip is caller-owned: record each turn yourself. Restoring is a replay through prefill, not a zero-cost snapshot — it costs a prefill of the history and counts against MaxNumTokens. Keep the system prompt in SystemMessage OR as a leading System message, not both (the native side prepends SystemMessage before the history). Takes precedence over HistoryJson when it is non-empty. Maps to the C API conversation_config_set_messages.

HistoryJson

Raw escape hatch for History: the messages as a JSON array string (the format Serialize(IEnumerable<LiteRtMessage>) produces). Use it to pass persisted history verbatim without re-parsing into typed messages, or to include content the typed model does not cover yet (e.g. image/audio parts). Validated as a JSON array when the conversation is created — throws ArgumentException otherwise. Ignored when History is non-empty.

public string? HistoryJson { get; init; }

Property Value

string

LoraPath

Path to a text LoRA adapter (weights file) to apply to this conversation. null (default) = no adapter. Requires a LoRA-enabled model, and the engine loaded with a matching LoraRank.

public string? LoraPath { get; init; }

Property Value

string

Remarks

The native side opens the file when the conversation is created, so a missing/unreadable path fails fast with LiteRtException at CreateConversation(LiteRtConversationOptions?) — not silently at first send. Maps to the C API session_config_set_lora_path. Not yet validated end-to-end in this binding (no LoRA adapter artifact was available at the time of writing): the plumbing is in place and a bad path is reported coherently, but a successful load against a real adapter has not been exercised here.

MaxOutputTokens

Maximum output tokens per response. 0 = engine default.

public int MaxOutputTokens { get; init; }

Property Value

int

PromptTemplate

Overrides the chat prompt template for this conversation with a custom Jinja template string. null (default) = the template the model file (or engine) provides. For advanced use — a template that does not match what the model was trained on degrades output quality.

public string? PromptTemplate { get; init; }

Property Value

string

Remarks

Maps to the C API conversation_config_set_prompt_template (native v0.15.0+). Inspect the result with RenderMessage(string) / RenderPreface() before relying on it.

Sampler

Optional sampler parameters. Null = engine default.

public LiteRtSamplerParams? Sampler { get; init; }

Property Value

LiteRtSamplerParams

StreamToolCalls

Stream the raw text of a tool call while the model is generating it. Default false: during a tool-call block the stream goes silent until the block completes and the parsed ToolCall chunk arrives whole. With this on, SendStreamingAsync(string, CancellationToken) additionally yields ToolCallDelta chunks — incremental, unparsed fragments of the tool-call text — as they are produced, followed by the usual complete ToolCall chunk. Use the deltas for progress display only; act on the final parsed chunk. Only meaningful when Tools are set.

public bool StreamToolCalls { get; init; }

Property Value

bool

Remarks

Maps to the C API conversation_config_set_stream_tool_calls (requires native LiteRT-LM v0.14.0+; older binaries throw EntryPointNotFoundException at conversation creation). The deltas do not affect the blocking Send(string) path.

SystemMessage

Optional system prompt applied to the conversation.

public string? SystemMessage { get; init; }

Property Value

string

ThinkingTokenBudget

Caps how many tokens the model may spend on its thinking block per reply: the decode loop tracks the model's thinking start/end token ids and cuts the block off at the budget. null (default) = no cap; -1 = explicitly infinite; 0 is treated by the native layer as "no budget" (NOT "no thinking"). Setting a budget with EnableThinking unset turns thinking ON (a budget only makes sense with a thinking block; the native thinking config defaults to enabled) — set EnableThinking = false explicitly if you want a budget configured but thinking off.

public int? ThinkingTokenBudget { get; init; }

Property Value

int?

Remarks

Unlike MaxOutputTokens — which caps the whole reply and starves the answer when thinking runs long — this bounds only the reasoning share. Maps to the C API thinking config (litert_lm_thinking_config_*, native v0.15.0+; older binaries throw EntryPointNotFoundException at conversation creation). The per-send ThinkingTokenBudget overrides this for one send.

Exceptions

ArgumentOutOfRangeException

The value is negative and not -1.

Tools

Tools the model may call (function calling). Null/empty = no tools.

public IReadOnlyList<LiteRtTool>? Tools { get; init; }

Property Value

IReadOnlyList<LiteRtTool>

VisualTokenBudget

Budget (in tokens) that image attachments may consume during prefill. 0 (default) = engine default. Lower it to cap how much of the context window an image eats on a vision model; only meaningful when sending image attachments. Applied per send via the C API conversation_optional_args_set_visual_token_budget.

public int VisualTokenBudget { get; init; }

Property Value

int