Table of Contents

Class LiteRtConversation

Namespace
LiteRtLmSharp
Assembly
LiteRtLmSharp.dll

A stateful conversation over a LiteRtEngine. Handles chat templating internally (mirrors the Gemini Chat APIs). Not thread-safe: serialize calls per instance — and serialize sends per engine, not just per conversation: multiple live conversations may be interleaved freely, but concurrent generation on different conversations of one engine is not supported by the native runtime (non-deterministic output and GPU-backend errors, upstream LiteRT-LM#2807; its thread-safety contract is undocumented).

public sealed class LiteRtConversation : IDisposable
Inheritance
LiteRtConversation
Implements
Inherited Members

Properties

IsContextFull

Whether this conversation's context is full: a completed send was detected to have run its reply into the KV overflow guard's ceiling (exact detection from the guard's own counts), or TokenCount has reached that ceiling (the engine's MaxNumTokens minus a reserve of ~128 tokens — the native executor prefills in fixed work groups of the model's prefill-signature lengths and rejects any send with fewer free entries than the smallest signature, so those trailing entries are unusable for a new send by construction). When true, the reply that got here was likely truncated by the guard's decode clamp (mid-sentence text, or a stream that just stopped), and the next send is guaranteed to throw LiteRtContextOverflowException — check this right after a send (or after a stream completes) to learn in the same turn that the conversation is over, instead of discovering it on the next call. Always false when the engine was loaded without an explicit MaxNumTokens (the limit is then internal to the engine and the guard is off).

public bool IsContextFull { get; }

Property Value

bool

TokenCount

Tokens currently held in this conversation's KV cache (prefill + decode, accumulated across turns). When this approaches the engine's MaxNumTokens the context is full and further generation degrades — manage history before that point. When the engine was loaded with an explicit MaxNumTokens, the send methods guard the limit and throw LiteRtContextOverflowException rather than let a send overflow the cache (which corrupts the native runtime).

public int TokenCount { get; }

Property Value

int

Methods

CancelProcess()

Cancels any in-flight generation on this conversation: a blocking Send(string) running on another thread, a SendAsync(string, CancellationToken), or an active SendStreamingAsync(string, CancellationToken) stream. This is the one operation that is safe to call from any thread while a send is in flight; when nothing is in flight it is a no-op. A cancelled blocking send fails with LiteRtException (the async variants translate that to OperationCanceledException). Afterwards, treat the conversation as consumed: dispose it and continue on a fresh one (restore prior turns via History if needed) — sending again on a cancelled conversation hangs inside the native runtime (LiteRT-LM v0.13.1, reproduced with Google's own binaries); the engine itself is unaffected. Mirrors the native cancel_process (the Kotlin binding's cancelProcess / the JS binding's cancel).

public void CancelProcess()

Clone()

Forks this conversation into a new, independent one that starts from a copy of the current prefilled (KV-cache) state — branch a conversation to explore several continuations without re-prefilling the shared prefix. The clone advances on its own; this conversation is untouched.

public LiteRtConversation Clone()

Returns

LiteRtConversation

Remarks

Call this only when the conversation is idle (no in-flight SendStreamingAsync(string, CancellationToken)) — conversations are not thread-safe. Dispose the clone like any conversation, before the engine. Cloning duplicates state in memory; to persist a conversation across process restarts use History instead.

The conversation must have advanced at least once (TokenCount > 0): the native clone duplicates the prefilled state, and creating a conversation prefills nothing — the system message, tools and History are only pushed into the KV cache by the first send. Cloning a never-sent conversation is not rejected natively but the result is unreliable: the first clone happens to work (it pays the parent's prefill itself), and after any other conversation runs on the engine the parent and its later clones continue that other conversation's context instead of their own (observed on LiteRT-LM v0.15.0 and v0.16.0, CPU and GPU). To keep a reusable "base" (system prompt + tools) to branch from, send one real turn on it first, then clone.

Exceptions

InvalidOperationException

The conversation has not advanced yet (TokenCount is 0), so there is no prefilled state to duplicate. Send a message on it first.

LiteRtException

The native clone failed. The usual cause is an engine/backend whose executor does not implement cloning (the native layer returns Unimplemented). The standard executors do — cloning is verified on both CPU and GPU (win-x64 WebGPU).

Dispose()

Disposes the conversation, freeing its native resources (and its config handles). Dispose a conversation, and any clones, before the LiteRtEngine it came from.

public void Dispose()

GetBenchmarkInfo()

Returns benchmark timings (prefill/decode tokens-per-second, time-to-first-token, init time) for this conversation, or null when benchmarking was not enabled (EnableBenchmark) or no turn has completed yet.

public LiteRtBenchmarkInfo? GetBenchmarkInfo()

Returns

LiteRtBenchmarkInfo

Remarks

The native per-turn accessors do not bounds-check their index, so this only reads a turn after confirming the corresponding turn count is > 0. Throws EntryPointNotFoundException on native binaries predating the benchmark API. Wraps a native surface Google still marks experimental in its own bindings; the reported values may change with the native runtime version.

RenderMessage(string)

Renders the user message text to the exact templated prompt string the model would receive, without sending it (conversation state is unchanged). Pair with Tokenize(string) to measure a turn's real token cost with the chat template included, or to inspect how a system prompt / history shape the rendered turn.

public string RenderMessage(string text)

Parameters

text string

Returns

string

Remarks

The rendered string is exactly what the next send would prefill: on a conversation whose preface (system message + tools + history) has not been consumed by a first send yet, it includes the preface; after the first send it is the turn alone — do not add RenderPreface() on top when measuring a first send's cost. Wraps a native entry point Google still marks experimental in its own bindings; the rendered format may change with the native runtime version.

RenderMessageRaw(string)

Low-level escape hatch: renders a raw message JSON to its templated prompt string. Use when you need full control over the wire format; otherwise use RenderMessage(string). Does not send or change conversation state.

public string RenderMessageRaw(string messageJson)

Parameters

messageJson string

Returns

string

RenderPreface()

Renders the conversation's preface — the templated preamble the model sees before the first user turn: the system message, any tools, and the restored History — to its exact prompt string, without sending (conversation state is unchanged). Pair with Tokenize(string) to measure how many tokens the system prompt / tools / history consume up front, or to inspect the templated preamble. Complements RenderMessage(string), which renders one user turn.

public string RenderPreface()

Returns

string

Remarks

Wraps a native entry point Google still marks experimental in its own bindings; the rendered format may change with the native runtime version. Requires native LiteRT-LM v0.14.0+ (throws EntryPointNotFoundException on older binaries).

Exceptions

LiteRtException

The native render call returned null.

Send(string)

Sends a user message and returns the reply (blocking). The answer text is Text; when the conversation has tools the model may instead return ToolCalls (see SendToolResults(IEnumerable<LiteRtToolResult>, LiteRtSendOptions?)), and reasoning models expose their thinking trace via Thinking. For an awaitable variant with mid-generation cancellation use SendAsync(string, CancellationToken); to cut a blocking send short from another thread, call CancelProcess().

public LiteRtResponse Send(string text)

Parameters

text string

Returns

LiteRtResponse

Send(string, IReadOnlyList<LiteRtAttachment>?, LiteRtSendOptions?)

Sends a user message with optional image/audio attachments (null/empty = text-only) and optional per-send options, returning the structured response. Attachments are appended after the text in content-part order and require the engine to have the matching modality enabled (see VisionBackend / AudioBackend) on a multimodal model.

public LiteRtResponse Send(string text, IReadOnlyList<LiteRtAttachment>? attachments, LiteRtSendOptions? options = null)

Parameters

text string
attachments IReadOnlyList<LiteRtAttachment>
options LiteRtSendOptions

Returns

LiteRtResponse

Exceptions

LiteRtContextOverflowException

Sending would overflow the KV cache sized by MaxNumTokens; the conversation is full and must be replaced. Only thrown when the engine was loaded with an explicit MaxNumTokens.

SendAsync(string, IReadOnlyList<LiteRtAttachment>?, LiteRtSendOptions?, CancellationToken)

Awaitable Send(string, IReadOnlyList<LiteRtAttachment>?, LiteRtSendOptions?) with true mid-generation cancellation — see SendAsync(string, CancellationToken) for the cancellation contract. attachments may be null/empty for a text-only send.

public Task<LiteRtResponse> SendAsync(string text, IReadOnlyList<LiteRtAttachment>? attachments, LiteRtSendOptions? options = null, CancellationToken cancellationToken = default)

Parameters

text string
attachments IReadOnlyList<LiteRtAttachment>
options LiteRtSendOptions
cancellationToken CancellationToken

Returns

Task<LiteRtResponse>

SendAsync(string, CancellationToken)

Awaitable Send(string) with true mid-generation cancellation: cancelling cancellationToken cancels the native inference (via CancelProcess()) and the task faults with OperationCanceledException. Treat a cancelled conversation as consumed: dispose it and continue on a fresh one (restore prior turns via History if needed) — sending again on it hangs inside the native runtime (LiteRT-LM v0.13.1, reproduced with Google's own binaries); the engine itself is unaffected. While the send is in flight, do not touch the conversation from other threads — cancellation is the only supported concurrent operation.

public Task<LiteRtResponse> SendAsync(string text, CancellationToken cancellationToken = default)

Parameters

text string
cancellationToken CancellationToken

Returns

Task<LiteRtResponse>

SendRaw(string, string?, LiteRtSendOptions?)

Low-level escape hatch: sends a raw message JSON and returns the raw response JSON. Use when you need full control over the wire format.

public string SendRaw(string messageJson, string? extraContext = null, LiteRtSendOptions? options = null)

Parameters

messageJson string
extraContext string
options LiteRtSendOptions

Returns

string

Exceptions

LiteRtContextOverflowException

Sending would overflow the KV cache sized by MaxNumTokens — the conversation is full (or the message's prefill leaves no room to reply); continuing would corrupt the native runtime. Only thrown when the engine was loaded with an explicit MaxNumTokens. A send carrying media or a non-null extraContext gets only the conversation-full check (its prefill cost is not measurable managed-side) — leave headroom under the limit for those.

SendStreamingAsync(string, IReadOnlyList<LiteRtAttachment>?, LiteRtSendOptions?, CancellationToken)

Streaming overload with optional image/audio attachments (null/empty = text-only) and per-send options. The attachments follow the text in content-part order and require the engine to have the matching modality enabled (VisionBackend / AudioBackend) on a multimodal model; otherwise the native send fails. The chunk kinds and cancellation behavior are identical to the text-only overload.

public IAsyncEnumerable<LiteRtStreamChunk> SendStreamingAsync(string text, IReadOnlyList<LiteRtAttachment>? attachments, LiteRtSendOptions? options = null, CancellationToken cancellationToken = default)

Parameters

text string
attachments IReadOnlyList<LiteRtAttachment>
options LiteRtSendOptions
cancellationToken CancellationToken

Returns

IAsyncEnumerable<LiteRtStreamChunk>

SendStreamingAsync(string, CancellationToken)

Sends a user message and streams the reply as LiteRtStreamChunk pieces, each tagged by Kind as answer text, reasoning ("thinking") text, or a tool call. Route on the kind rather than assuming a global order; concatenate same-kind text deltas to rebuild the answer and the thinking trace. With EnableThinking on, reasoning models emit the thinking trace before the answer. A ToolCall chunk only appears when the conversation was created with tools — handle it like the blocking Send(string) loop (run the tools, then SendToolResults(IEnumerable<LiteRtToolResult>, LiteRtSendOptions?)). With StreamToolCalls on, raw ToolCallDelta progress fragments additionally precede that complete tool-call chunk.

public IAsyncEnumerable<LiteRtStreamChunk> SendStreamingAsync(string text, CancellationToken cancellationToken = default)

Parameters

text string
cancellationToken CancellationToken

Returns

IAsyncEnumerable<LiteRtStreamChunk>

SendToolResults(IEnumerable<LiteRtToolResult>, LiteRtSendOptions?)

Sends the results of executed tools back to the model and returns its next response. Call after a LiteRtResponse with IsToolCall = true.

public LiteRtResponse SendToolResults(IEnumerable<LiteRtToolResult> results, LiteRtSendOptions? options = null)

Parameters

results IEnumerable<LiteRtToolResult>
options LiteRtSendOptions

Returns

LiteRtResponse

Exceptions

LiteRtContextOverflowException

Sending would overflow the KV cache sized by MaxNumTokens — the guard that keeps a long tool loop from crashing the native runtime. Only thrown when the engine was loaded with an explicit MaxNumTokens.

SendToolResultsAsync(IEnumerable<LiteRtToolResult>, LiteRtSendOptions?, CancellationToken)

Awaitable SendToolResults(IEnumerable<LiteRtToolResult>, LiteRtSendOptions?) with true mid-generation cancellation — see SendAsync(string, CancellationToken) for the cancellation contract.

public Task<LiteRtResponse> SendToolResultsAsync(IEnumerable<LiteRtToolResult> results, LiteRtSendOptions? options = null, CancellationToken cancellationToken = default)

Parameters

results IEnumerable<LiteRtToolResult>
options LiteRtSendOptions
cancellationToken CancellationToken

Returns

Task<LiteRtResponse>