Class LiteRtContextOverflowException
- Namespace
- LiteRtLmSharp
- Assembly
- LiteRtLmSharp.dll
Raised instead of letting a send overflow the conversation's KV cache. The KV cache is sized by
MaxNumTokens and the native runtime does not police it: a send
that grows past the limit writes beyond the allocated cache, corrupting native memory — the typical
symptom is a deferred process crash (0xC0000005 / 0xC0000374) on a LATER call, with
nothing pointing back at the overflowing send. The binding throws this instead, before native
state is damaged.
public sealed class LiteRtContextOverflowException : LiteRtException, ISerializable
- Inheritance
-
LiteRtContextOverflowException
- Implements
- Inherited Members
Remarks
Thrown by the send methods of LiteRtConversation when the engine was loaded
with an explicit MaxNumTokens and either (a) the conversation is
already at the limit (TokenCount ≥ MaxNumTokens), or (b) the message's
measured prefill cost leaves no room to decode a reply. Without an explicit
MaxNumTokens the limit is internal to the engine and the guard is off.
A conversation that hit the limit cannot continue: dispose it and start a fresh one — with a
trimmed History if the thread must go on — or reload the
engine with a larger MaxNumTokens. In the Extensions.AI stateful
mode the client evicts the live conversation automatically when this throws; resuming its
ConversationId then fails with ArgumentException like any evicted id.
Constructors
LiteRtContextOverflowException(string, int, int)
Creates the exception. tokenCount is the conversation's KV-cache
size when the send was rejected; maxNumTokens is the engine's configured
context limit.
public LiteRtContextOverflowException(string message, int tokenCount, int maxNumTokens)
Parameters
Properties
MaxNumTokens
The engine's configured context limit (MaxNumTokens).
public int MaxNumTokens { get; }
Property Value
TokenCount
The conversation's KV-cache token count when the send was rejected (TokenCount).
public int TokenCount { get; }