LiteRT-LM C ABI — reference for the .NET binding
Source of truth:
c/engine.hin the official repo. This document summarizes the ABI as verified against the real binary.
Current state (v0.16.0, Google's official C API prebuilts — see native-build.md): the findings about the community binary
0.12.0-aand the interim commit032334d8(conversation_config_createcrash, missingget_token_count, blockingsend_messagereturning null, streaming segfault) are HISTORICAL — first resolved by compiling our own binaries from thev0.13.1tag with the matching header (the self-built era, v0.13.1 → v0.15.0), and since v0.16.0 by shipping upstream's official prebuilts (144litert_lm_*exports, 117 bound). Today config/system-prompt/sampler, tools, streaming and token count work on all 5 platforms. Speculative decoding, the benchmark API, and the engine cache-dir setting were bound on 2026-06-15; multimodal image/audio messages on 2026-06-17; the tokenizer surface (tokenize/detokenize + start/stop tokens) on 2026-06-19; and the v0.14.0 surface (LoRA, CPU thread counts, per-send output cap, tool-call streaming, preface rendering, plus the internal sampler-builder migration) on 2026-07-10 (seeroadmap.mdfor the C-API coverage count, now 84/109, the v0.14.0 ABI changes section below, and the multimodal section below). The notes below are kept as a diagnostic record.Caveat (2026-06-12): "all 5 platforms" for tools was validated by hand on win-x64/Android; on desktop Linux the tools + constrained-decoding path had never actually run in CI (regular CI has no model), and a Linux Mint user reported the process dying silently with tools enabled — matches upstream LiteRT-LM#2149 (C-API shared lib segfaults/hangs in decode on Ubuntu 24.04; static CLI works). The
model-tests.ymlworkflow now exercises streaming + tools on linux-x64 with gemma-4-E2B-it (on each push, via ci.yml).
Viability summary (verified)
- The
c/engine.hheader declares a flat C API (extern "C") with opaque pointers — ideal for P/Invoke. ~89 functions,litert_lm_prefix. - Exported on Windows via
__declspec(dllexport); on Linux/macOS viavisibility("default"). - Verified against the binary:
LiteRtLm.dll(flutter_gemma prebuilt, tagnative-v0.12.0-a) exports 89litert_lm_*functions in its export table (confirmed withdumpbin /exports). → P/Invoke against these binaries was viable from day one, without building from source.
Origin of the prebuilt binaries (PoC phase)
flutter_gemma publishes LiteRT-LM natives as GitHub Release assets on its own repo:
- Base:
https://github.com/DenisovAV/flutter_gemma/releases/download/native-v<version>/ - Version used during the PoC:
0.12.0-a(tagnative-v0.12.0-a). - Relevant desktop assets:
litertlm-windows_x86_64.tar.gz(sha256b7264091c05001ef84e53761dfee331f761e3a2362b36b28ab2ce39666400d76)litertlm-linux_x86_64.tar.gz(sha256930296b010ecc316c6b6fc4ed1c722b275b4064b59b5aad8ff7b858e9149c0d7)
- Main lib:
LiteRtLm.dll(Win) /libLiteRtLm.so(Linux) → P/Invoke name:LiteRtLm.
Required companions (Windows x64)
LiteRtLm.dll resolves the lib-prefixed copies via PE imports (at LoadLibrary time):
libLiteRt.dll, libGemmaModelConstraintProvider.dll, libLiteRtTopKWebGpuSampler.dll,
libLiteRtWebGpuAccelerator.dll, plus the DXC runtime (dxcompiler.dll, dxil.dll) and the
optional Intel NPU set (LiteRtDispatch.dll, openvino*.dll, tbb*.dll). Every .dll in
the tarball must ship together in the output directory. The CPU backend works without the
NPU part.
License note: those binaries are Apache-2.0 (LiteRT-LM) repackaged by flutter_gemma. For production we build our own (see docs/native-build.md) and/or will consume the official target from #2154.
Minimal flow (Conversation API — high level, recommended)
Handles chat templates internally; mirrors the Gemini Chat APIs via JSON.
// 1. Settings
LiteRtLmEngineSettings* s = litert_lm_engine_settings_create(model_path, "cpu", NULL, NULL);
litert_lm_engine_settings_set_max_num_tokens(s, 512); // optional
// 2. Engine (heavy, owns the weights)
LiteRtLmEngine* e = litert_lm_engine_create(s);
// 3. Conversation (NULL config = defaults)
LiteRtLmConversation* c = litert_lm_conversation_create(e, NULL);
// 4a. Blocking send
LiteRtLmJsonResponse* r = litert_lm_conversation_send_message(c, msg_json, NULL, NULL);
const char* out_json = litert_lm_json_response_get_string(r); // string owned by r
// 4b. ...or streaming (callback on a background thread)
litert_lm_conversation_send_message_stream(c, msg_json, NULL, NULL, cb, user_data);
// 5. Release in reverse order
litert_lm_json_response_delete(r);
litert_lm_conversation_delete(c);
litert_lm_engine_delete(e);
litert_lm_engine_settings_delete(s);
JSON contract (verified in c/engine_test.cc)
- User message (
message_json):{"role": "user", "content": [{"type": "text", "text": "Hello"}]} - Response (
litert_lm_json_response_get_string): same shape; the text lives atresponse["content"][0]["text"]. - System message (for
litert_lm_conversation_config_set_system_message, content is an object, not an array):{"type":"text","text":"You are a helpful assistant."}
Streaming callback
typedef void (*LiteRtLmStreamCallback)(void* callback_data, const char* chunk,
bool is_final, const char* error_msg);
chunk: text fragment (valid only during the call → copy it).error_msg: NULL on success.is_final: true on the last chunk → signal completion. Invoked from a background thread.
.NET marshalling conventions
const char*strings = UTF-8 →StringMarshalling.Utf8in[LibraryImport].- C
bool= 1-byte bool →[MarshalAs(UnmanagedType.U1)]/byte. - Opaque pointers → one
SafeHandleper type; release with its*_delete. - x64 has a single calling convention; declare Cdecl explicitly
(
[UnmanagedCallConv(Cdecl)]) for portability. - Callback: use
[UnmanagedCallersOnly(Cdecl)]+ aGCHandleincallback_data(AOT-friendly, no delegate marshalling).
Key header types/structs
LiteRtLmSamplerParams { LiteRtLmSamplerType type; int32 top_k; float top_p; float temperature; int32 seed; }(pre-v0.14.0 shape; v0.14.0 removed this by-value struct in favor of an opaque builder, see v0.14.0 ABI changes below).LiteRtLmSamplerType: 0 Unspecified, 1 TopK, 2 TopP, 3 Greedy. v0.14.0 dropped theUnspecifiedmember (see below); the publicLiteRtSamplerType.Unspecifiedis retained at0.LiteRtLmInputData { LiteRtLmInputDataType type; const void* data; size_t size; }(multimodal; text=UTF-8).LiteRtLmInputDataType: Text, Image, ImageEnd, Audio, AudioEnd.
v0.14.0 ABI changes (sampler struct → opaque builder)
The binding is pinned to LiteRT-LM v0.14.0 (repinned from v0.13.1 on 2026-07-10). v0.14.0 grew the C API from 89 to 109 functions (84 now bound). The one breaking ABI change the binding depends on:
- Sampler params: by-value struct → opaque builder. v0.14.0 removed the by-value
LiteRtLmSamplerParamsstruct (the pre-v0.14.0 shape above) and replaced it with an opaque builder:sampler_params_create+set_top_k/set_top_p/set_temperature/set_seed, copied into the session config and deleted immediately. The binding was rewired to the builder (required migration; the by-value struct is gone). The native enum also dropped itsUnspecifiedmember; the publicLiteRtSamplerType.Unspecifiedis retained (value0) and now sends no sampler params at all (the executor's internal default), which is the same effective behavior as v0.13.1, where the unspecified type made the native sampler factory defer to the executor. (The old doc claim that the engine resolvedUnspecifiedtoTopPwas false.)
Version match is mandatory. Because the binding now calls the v0.14.0 builder functions (and the other
16 newly bound entry points), it requires version-matched v0.14.0+ native binaries. Run it against
older natives (v0.13.1 or earlier) and the first call into a missing function throws
EntryPointNotFoundException, the same skew failure mode documented for get_token_count on
0.12.0-a above. Always restore the native-v<tag> release that matches the managed package
(the LiteRtLmSharp.runtime.<rid> package version); never mix versions.
Empirical findings on the prebuilt native-v0.12.0-a binary (VERIFIED at runtime)
Tested with gemma-4-E2B-it.litertlm (CPU/XNNPACK) from .NET:
- Generation works end-to-end (blocking and streaming). Engine loads in ~0.2 s (mmap).
- Streaming chunks are full JSON objects per token, not plain text:
{"role":"assistant","content":[{"type":"text","text":"1"}]}. → every chunk must be parsed forcontent[0].text(not just the final blocking-path response). litert_lm_conversation_config_createtriggers an AccessViolation (0xC0000005) in this binary, despite being in the export table (ordinal 28). It is version skew: the header came frommain(~0.13+) while the binary was 0.12.0-a. → Workaround at the time: create conversations with a NULL config (litert_lm_conversation_create(engine, NULL)).- Blocking
litert_lm_conversation_send_messagereturned NULL in some conditions where streaming worked. → The streaming path was the robust one in this binary. litert_lm_conversation_get_token_countis NOT in this binary (throwsEntryPointNotFoundException); it was added upstream after 0.12.0-a.MaxNumTokensis the TOTAL context window (KV cache = prompt + response, accumulated across turns). If small (e.g. 1024) a long answer fills it and later turns overflow and degenerate into incoherent text (observed symptom: answer cut mid-word, then garbage like "Laptop"). Raising it to 4096 fixed a coherent multi-turn chat. Not a binding bug; it's LLM context management. For production: expose/manage history and, when the binary allows it, cap per-turn output viasession_config_set_max_output_tokens.
Sync lesson: the header and the binary must come from the same LiteRT-LM tag. The skew explains (3) and (4); building from a pinned tag eliminates it.
Tool / function calling (verified wire format)
The C API exposes tools via the conversation config (requires a skew-free binary):
litert_lm_conversation_config_set_tools(config, tools_json) +
litert_lm_conversation_config_set_enable_constrained_decoding(config, true).
- Tool definition (OpenAI/Gemini FunctionDeclaration style, from
c/engine_test.cc):[{"type":"function","function":{"name":"get_current_weather","description":"...", "parameters":{"type":"object","properties":{"location":{"type":"string"}},"required":["location"]}}}] - Tool-call response (what
send_messagereturns; Gemma 4 / FunctionGemma docs):{"role":"assistant","tool_calls":[{"type":"function", "function":{"name":"get_current_weather","arguments":{"location":"Tokyo"}}}]} - Tool-result message (sent back via
send_message):{"role":"tool","content":[{"name":"get_current_weather","response":{"temperature":15}}]}
Wrapper surface: LiteRtTool, LiteRtConversationOptions.Tools + EnableConstrainedDecoding,
conv.Send(text) → LiteRtResponse (.Text or .ToolCalls), conv.SendToolResults(...), and
conv.SendRaw(json) as an escape hatch. The parser is tolerant (function.arguments as
object or string; fallback to top-level name/args) and always exposes RawJson.
VALIDATED end-to-end with our own binary + gemma-4-E2B-it (CPU): define tool → model
emits tool-call → execute → re-inject → correct final answer. conversation_config_create
no longer crashes (matched header+binary).
Gemma template quirk: with constrained decoding the arguments arrive with
<|"|>tokens as quotes (<|"|>Tokyo<|"|>). The parser sanitizes them (StripControlTokens/CleanJson) →"Tokyo".
Multimodal messages (image / audio) — verified wire format
Multimodal works on the high-level Conversation API — no need for the low-level Session/InputData path. Two layers:
- Engine — enable the encoders at
engine_settings_create(model, backend, vision_backend, audio_backend). The two trailing args areconst char*:"cpu"/"gpu"to enable that modality,NULLto leave it off (the documented sentinel — pass a real C#null, not""). Confirmed by upstreamengine_test.ccCreateSettingsWithVisionAndAudioBackend(vision="gpu",audio="cpu"). A model can constrain its audio backend. Gemma 4's audio sub-model requires CPU:audio_backend="gpu"makesengine_createfail withINVALID_ARGUMENT: Audio backend constraint mismatch. Model requires one of [cpu] but Audio backend is GPU— on any platform, not a win-x64 quirk (verified 2026-06-17; the model-tests macOS GPU leg confirms the same skip). Run audio on CPU for such models even when the main backend is GPU; the vision encoder is unconstrained and runs on GPU. (Upstream's own test pairsvision="gpu"withaudio="cpu".)engine_settings_set_max_num_images(settings, int)exists but the header says it is legacy-only (the current engine path ignores it) — bound for completeness; the real per-turn knob is the visual token budget below. - Message — attach media as extra content parts in the same
message_jsonthe text path uses:
Image and audio parts are interchangeable in two forms (byte-verified against{"role":"user","content":[ {"type":"text","text":"Describe this image: "}, {"type":"image","blob":"<BASE64 image bytes>"} ]}runtime/conversation/model_data_processor/data_utils.ccLoadItemData):{"type":"image"|"audio","blob":"<base64>"}—blobis base64 (decoded withabsl::Base64Unescape; a bad string →InvalidArgumentError("Failed to decode base64 blob.")). The .NET side mustConvert.ToBase64String(bytes).{"type":"image"|"audio","path":"/abs/path"}— memory-mapped natively (MemoryMappedFile::Create), no base64 round-trip. Desktop only; the path must be readable by the native process.- No
mime_type/mime/data/image_urlfield exists or is read — the media kind is the"type"string alone. Content-part order is preserved and significant (interleave text/media as seen).
- Visual token budget — optional per-send override created via
conversation_optional_args_create()→conversation_optional_args_set_visual_token_budget(args, int)→ passed as the last arg ofsend_message/send_message_stream→conversation_optional_args_delete(args). For streaming the args must outlive the whole stream (the native decode thread reads them during prefill), so the binding frees them in the iterator'sfinally, alongside the callbackGCHandle.
The v0.13.1
c/engine_test.ccexercises only text messages through the Conversation API (andset_visual_token_budget(args, 100)with a text-only message), so the image/audio JSON shape is taken from the parser source (data_utils.cc) and the model-specific data processors (gemma4_data_processor.ccuses an image and an audio preprocessor), not from the test file.
Wrapper surface: LiteRtAttachment.Image/ImageFile/Audio/AudioFile,
LiteRtConversation.Send(text, attachments) /
SendStreamingAsync(text, attachments, ct), LiteRtEngineOptions.VisionBackend/AudioBackend/
MaxNumImages, LiteRtConversationOptions.VisualTokenBudget.
The vision/audio executor needs a session config on the conversation. The encoder executor only loads (
TryLoadingVisionExecutor) when the conversation was created with a session config; without one a first image/audio send fails withINVALID_ARGUMENT: Vision executor should not be null, please TryLoadingVisionExecutor() first. A bare session config (nothing set on it) is enough — verified by driving the C API directly. The binding now attaches one automatically whenever the engine was loaded withVisionBackend/AudioBackend(LiteRtConversation.Createis passed anengineIsMultimodalflag), so a plainCreateConversation()can send attachments. SettingMaxOutputTokensor aSampleralso creates a session config, which is why the model-backed tests (which setMaxOutputTokens = 16) never hit this.Earlier mis-diagnosis (corrected 2026-06-20). The 2026-06-17 smoke test blamed
MaxNumTokens("use ≥ 4096; 2048 fails"). That was a confound — the failing case used a bare conversation (no session config) while the working case used the sample'sMaxOutputTokenscap. A 2×2 over{2048,4096} × {default, capped}showedMaxNumTokensis irrelevant: 2048 works with a session config, 4096 fails without one. The realMaxNumTokensfloor is just enough to hold the media's tokens: for a 64×64 image (~283 tokens) it fails atMaxNumTokens = 256and works at384.Wrapped error. When a send still fails this way (e.g. the model is not multimodal, or
VisionBackend/AudioBackendwas left unset, orMaxNumTokenscannot hold the media), the binding throws a managedLiteRtExceptionnaming those causes, on both the blocking and streaming send paths. It flags the message as media-bearing by parsing the content parts only when a send fails, so a normal send pays nothing.
VALIDATED with our own win-x64 binary + gemma-4-E2B-it (2026-06-17): the self-built lib links the
vision/audio executors. CPU: a red PNG expanded to a ~261-token vision block (text-only prefill 28 →
with-image 289) and the model answered "…a solid, vibrant red color"; a real spoken 5→0 countdown
(countdown.mp3) added ~130 audio tokens (35 → 165) and the model transcribed "Five, four, three, two,
one, zero." win-x64 GPU: vision runs on the WebGPU/D3D12 backend (same red result, 28 → 289); audio
runs on CPU (35 → 164, "5 4 3 2 1 0") because the model's audio sub-model is CPU-constrained
(audio_backend="gpu" → "Audio backend constraint mismatch. Model requires one of [cpu]"), which is a
model property, not a platform one (the macOS GPU leg confirms the same). The model-backed tests assert
the token delta (deterministic proof the encoder ran) and log the transcription. (An earlier synthetic
sine tone yielded a canned "I cannot process audio"; a real clip transcribes correctly, so the fixture is
an embedded countdown.mp3.)
Streaming: regression in 032334d8, RESOLVED in v0.13.1
SendMessageStreamingAsync segfaulted (exit 139) with the interim-commit binary 032334d8 —
the native decode thread crashed BEFORE the first callback (rc=0; [cb:enter] never reached).
It was not managed code (worked on 0.12.0-a), not the WebGPU sampler (#2073), and not
litert_link_capi_so (present in both builds): it was a regression in that commit.
Building from the v0.13.1 release tag fixes it. Verified: streaming OK, tools OK,
get_token_count now exported (89 funcs), and the streaming→tools sequence in one process
(which used to segfault) passes. Test suite 4/4 on v0.13.1.
Lesson: pin to a release tag, never an arbitrary commit (more stable, and it is the sync target with Google).
native-release.ymldefaults to the current pin (v0.16.0).
Tokenizer (tokenize / detokenize / start-stop tokens) — verified
The engine exposes the model's own tokenizer, so exact token counts can be computed without running
inference (16 functions, all bound). The shape is opaque-pointer + getters, freed by the matching
*_delete; the const int* / const char* getters point INTO the owning object, so the .NET side
copies the data out before disposing the handle.
LiteRtLmTokenizeResult* t = litert_lm_engine_tokenize(engine, "Hello"); // NULL on failure
size_t n = litert_lm_tokenize_result_get_num_tokens(t);
const int* ids = litert_lm_tokenize_result_get_tokens(t); // valid while t lives
litert_lm_tokenize_result_delete(t);
LiteRtLmDetokenizeResult* d = litert_lm_engine_detokenize(engine, ids, n); // inverse
const char* text = litert_lm_detokenize_result_get_string(d); // UTF-8, owned by d
litert_lm_detokenize_result_delete(d);
- Start (BOS) / stop (EOS) tokens.
litert_lm_engine_get_start_tokenreturns aLiteRtLmTokenUnion*(or NULL);litert_lm_engine_get_stop_tokensreturns aLiteRtLmTokenUnions*collection. ATokenUnionis a tagged value:litert_lm_token_union_get_typeiskLiteRtLmTokenUnionTypeString(read..._get_string) orkLiteRtLmTokenUnionTypeIds(read..._get_ids(out_tokens, out_num)→ returns 0 on success). Ownership trap:litert_lm_token_unions_get_token_atreturns a NEW union the caller mustlitert_lm_token_union_delete— unlike theconst-view getters above, it is not owned by the collection. - .NET surface:
LiteRtEngine.Tokenize(string) → int[],Detokenize(ReadOnlySpan<int>) → string,GetStartToken() → LiteRtTokenUnion?,GetStopTokens() → IReadOnlyList<LiteRtTokenUnion>(LiteRtTokenUnion.Kind/Text/Ids). VALIDATED on win-x64 CPU withgemma-4-E2B-it(2026-06-19): text round-trips through ids, counts are deterministic and monotone, and the model reports a non-empty set of stop tokens (start = id[2], stops = ids[1]/[50]/[106]). - Detokenize returns the tokenizer's surface form, not post-processed text. For Gemma's SentencePiece
tokenizer,
litert_lm_engine_detokenizejoins the pieces with the ▁ (U+2581) space meta-symbol — e.g. ids for "The quick brown fox." come back asThe▁quick▁brown▁fox.. The chat/response path renders ▁ as spaces; the raw tokenizer surface does not. The binding passes the native string through verbatim (post-processing it would be wrong for non-SentencePiece tokenizers and would corrupt literal ▁).
Desktop GPU backend (WebGPU) — expected behavior
- The
"gpu"backend on desktop uses native WebGPU (Dawn), NOT a browser. It is a portable GPU layer mapping to: Direct3D 12 on Windows, Vulkan on Linux, Metal on macOS. (On Android: OpenCL/Vulkan.) - Verified: with
Backend="gpu"the log selects the discrete GPU (e.g.NVIDIA RTX 3080, backend=Direct3D 12) and runs the transformer layers on GPU (delegate_webgpu.cc,delegate_kernel.cc). Enabling companions (already shipped):libLiteRtWebGpuAccelerator.dll(loaded at runtime by base name from libLiteRt) +dxcompiler.dll/dxil.dll(DirectX Shader Compiler, loaded lazily by Dawn at the first shader compile). - Seeing "Created TensorFlow Lite XNNPACK delegate for CPU" alongside is normal: non-GPU ops + mmap'd embeddings run on CPU (mixed delegation). The bulk (matmuls) runs on GPU. Init is slower than CPU (~1.6 s vs ~0.2 s) due to weight upload and kernel compilation.
GPU sampler falls back to CPU on Windows/macOS — upstream bug (#2073), NOT the binding
Historical (self-built era, v0.13.1 → v0.15.0). The official v0.16.0 monolith embeds the TopK samplers, so this fallback no longer occurs; an explicit TopK sampler works on desktop GPU.
- Symptom:
Could not load symbol LiteRtTopKWebGpuSampler_UpdateConfig→Falling back to CPU sampling. - Cause (verified with
dumpbin /exports): the WindowsLiteRtTopKWebGpuSampler.dllexports only 3 of 7 functions (_Create,_Destroy,_SampleToIdAndScoreBuffer); missing_UpdateConfigetc. That is issue #2073 (Linux/Android ship all 7). - The fallback message mentions
.so/LD_LIBRARY_PATH/prebuilt/: it is a non-localized, Linux-centric log string; on Windows the equivalent file is the.dllwe already ship. It does not actually try to load a.so. - Impact: sampling (token selection) runs on CPU; the matmuls stay on GPU. Sampling is tiny compared to the matmuls → negligible performance impact. Output is correct either way.
- Definitive fix: our own build or a new prebuilt once #2073 is resolved upstream.
Note:
I0000 …logs show up despiteSetMinLogLevel(3)because they are emitted beforeabsl::InitializeLog()(straight to STDERR); our log level cannot silence them.
Official shared-library status
- Since v0.16.0 upstream publishes official C API prebuilts on every release
(
litert_lm_c_api-<version>.zipfor linux/windows/macos/android,CLiteRTLM.xcframework.zipfor iOS), and 1.2.0 ships those — see native-build.md. The self-built target (native/patch_c_api.sh, v0.13.1 → v0.15.0) is retired. - Before that: as of v0.14.0 upstream had its own
cc_binary litert-lminc/BUILD(the Python-wheel build); earlier tags shipped only the Bazelcc_library(:engine,:engine_cpu) andadd_litertlm_library(... STATIC)in CMake, with no shared-lib target (issue #2154 / PR #2155). The PoC consumed flutter_gemma'sLiteRtLm.dll/.so.