Technology & Research · · Team LATENT
Can Liquid AI d1 run on the same PC as your game?
d1-3B and the experimental 600M return decisions in one pass. What the published 8 ms really measures, and how to budget memory and integrate decisions without stalling a game.
An enemy NPC needs to choose whether to pursue, hide or call for help. That decision need not involve generating a long answer. Liquid AI released d1-3B and the experimental d1-omni-600M on October 7, 2026: models that return typed decisions for questions supplied by an application.
Integration paths exist. Running smoothly on the same PC as a game is a separate question, and the published 8 ms on an RTX 4090 does not settle it. This article examines public information as of October 9, 2026. LATENT has not run the models or measured them alongside a game.
Liquid AI release announcement
Key takeaways
- d1 returns choices or scores in one forward pass without generating prose. It can be evaluated for NPC action selection or intent classification; the game still executes movement and attacks.
- The published RTX 4090 result of 8 ms uses a warm, optimized short-question workload. It does not establish latency for long inputs or competing rendering workloads; headroom on the target PC needs separate measurement.
- Official GGUF releases provide a llama.cpp route and smaller quantized files. File size is not runtime memory usage, and the weights use lfm1.0 with commercial conditions.
A model for NPC decisions
d1-3B accepts text and images. The early research 600M release accepts text with images or text with audio, but not images and audio together. Its audio training uses requests between an English speaker and an assistant, with clips cut at 30 seconds. Generalization to Japanese in-game speech needs evaluation.
| Question type | Example game design |
|---|---|
| choice | Choose among allowed actions such as retreat, wait and call for help. |
| noul | Return a probability that an utterance asks for help, without composing a reply. |
| score | Rate an ordered scale such as calm, alert and critical. |
3B input and output specification
600M input and audio limitations
The shared idea with our Strands Decider article is to return a selection without text generation and leave execution to the application. Here, the interesting extension is to images and, for 600M, audio. The backbones also differ: d1-3B builds on LFM2.5-VL-3B, while 600M uses a bidirectional encoder. Comparing an older Decider timing with these results on different hardware and inputs would not establish a winner.
Related: how Strands Decider works
The model selects; the game acts
An illustrative NPC design, not a tested integration
1 Select state
Extract health, distance and the recent utterance needed for the decision.
2 Score options
Send allowed actions to d1. Provide movement coordinates and dialogue separately.
3 Validate rules
Check that the action remains valid, then pass it to the behavior tree or state machine.
What 8 ms measures and why input size matters
These are developer measurements, not concurrent gaming tests. The RTX 4090 row uses BF16 and the median of 20 warm runs. Single questions use CUDA graphs through model.compile; the reported uncompiled result is 16 ms. A new input shape incurs initial kernel selection or compilation work.
| RTX 4090 workload | Published latency |
|---|---|
| Short single question | 8 ms |
| Three questions over one state | 21 ms |
| 3.4K-token state | 102 ms |
| 384 px image | 17 ms |
Published results and conditions: d1-3B model card
One pass still has to process its input. At 102 ms, the 3.4K-token workload takes about 12.8 times the 8 ms result. Likewise, throughput of 475 states per second with 64 packed states does not mean an isolated request returns in about 2 ms. Batch throughput and response latency differ. No speed results accompany this 600M release, so smaller size alone does not establish that it is faster.
A frame at 60 fps lasts about 16.7 ms, but rendering and game logic already need part of that time. Subtracting 8 ms from 16.7 ms does not establish a concurrent budget. Inference and rendering share GPU compute, memory bandwidth, power limits and queues. Average fps can stay acceptable while some individual frames become much slower.
Separate weight size from runtime budgets
Start with a weights-only estimate. Using the reported 3.12 billion and 587 million parameters, multiply parameter count by bytes per value. The table is neither measured VRAM usage nor a recommended RAM capacity. GB means one billion bytes; GiB means 2 to the power of 30 bytes.
| Estimated full weights | 16bit | 32bit |
|---|---|---|
| d1-3B | 6.24 GB / 5.81 GiB | 12.48 GB / 11.62 GiB |
| d1-omni-600M | 1.174 GB / 1.09 GiB | 2.348 GB / 2.19 GiB |
Separately, the official GGUF releases list a 1.67 GB 3B Q4_K_M main file and a 407 MB 600M Q8_0 main file. These are distribution sizes, not all-modality totals or peak runtime usage. Media projectors, input processing, intermediate computations and runtime allocations add to them. Quantization format and CPU/GPU placement also matter.
Official 3B GGUF and distribution sizes
Official 600M GGUF and distribution sizes
| Distribution combination | Sum of published file sizes |
|---|---|
| 3B Q4_K_M main file plus Q8 image projector | 1.67 GB + 583 MB ≈ 2.25 GB |
| 600M Q8_0 main file plus Q8 media projector | 407 MB + 263 MB ≈ 670 MB |
| 600M F16 main file plus F16 media projector | 764 MB + 441 MB ≈ 1.205 GB |
These totals add the displayed distribution sizes. They exclude runtime placement and working allocations, so they are not minimum free-RAM or free-VRAM specifications.
Budget three resources on one PC
Fitting in memory and meeting a deadline are separate conditions
GPU memory
Game textures and render buffers plus model weights and temporary allocations.
GPU time
Rendering plus inference. Asynchronous work still shares compute and bandwidth.
CPU and RAM
Physics, audio and input processing plus CPU inference or model loading.
Suppose a game uses 6 GiB of an 8 GiB GPU, leaving 2 GiB. The estimated 5.81 GiB of full 16-bit 3B weights would not fit there. But a 1.67 GB quantized file being smaller than 2 GiB still does not prove it will fit at runtime. Moving work to the CPU can reduce GPU demand, while introducing RAM, CPU-time and transfer constraints. This is a budgeting example, not a measurement of a particular game or PC.
Keep inference off the rendering path
A sensible first target is a strategy change every few seconds or a decision at a dialogue boundary, rather than per-frame aiming or collision detection. Submit a request when health falls, an enemy is lost or the player finishes speaking. There is no need to turn an already deterministic distance check into a language-model request.
In an illustrative design, rendering continues with the current behavior while a separate inference process receives a state snapshot. Tag the request with NPC ID, state version and deadline; discard stale results. Keep the existing behavior tree running if a reply is late. Limiting each NPC to one pending request helps prevent a queue of decisions about obsolete situations.
Frequency matters. Twenty NPCs making two requests each per second produce 40 requests per second. Even at a hypothetical serial 8 ms per request, that adds up to 320 ms of processing per second. This is not measured GPU utilization under contention. If requests take 100 ms, the same serial worker cannot keep up. A design can reserve frequent decisions for important visible NPCs and use conventional logic for distant ones.
Multiple questions can share a state, but avoid mixing knowledge that different NPCs should not share. When the game can supply a compact internal state, its input and processing are easier to control than capturing and interpreting the screen again. For voice input, measure the player-visible delay including recording and preprocessing.
Python and llama.cpp integration paths
Official Python examples use Transformers and PyTorch. The requirements are Transformers 5.14 or later for 3B and 5.15 or later for 600M. Examples use float32 on CPU, bfloat16 for 3B on GPU and float16 for 600M on GPU. Do not apply the same dtype indiscriminately: the 600M card warns against bfloat16.
Loading requires trust_remote_code=True, which executes supplied model code. Review and pin the code and revision before adopting it. This investigation did not install dependencies or download weights.
The models are not Python-only. The official GGUF READMEs show loading them into llama.cpp llama-server and sending state and questions to /v1/systemone. A 600M question must fit in one batch; the example sets -b 4096 -ub 4096. Images and audio also require the appropriate projectors.
A local HTTP connection from Unity or Unreal is a possible integration design, not evidence of an official engine plugin. llama.cpp merged 3B support in 88dcc46 on October 7 and 600M support in a657f7e on October 8. Use a build containing the relevant change and verify the target OS, dedicated endpoint and required modalities. The minimum packaged release number and Windows runtime behavior have not been tested here.
Do not infer that ordinary Ollama or LM Studio chat interfaces expose this decision API just from generic Hugging Face usage suggestions. Check the dedicated call documented in the author-maintained README, rather than a text-generation chat interface.
Calling 600M through llama.cpp
Commercial conditions for shipping a game
The model cards specify lfm1.0, the LFM Open License v1.0. Although based on Apache 2.0, it is not Apache or MIT. Commercial use has a US$10 million annual-revenue threshold; commercial use by a legal entity exceeding the threshold is not licensed under this agreement. Do not assess eligibility using only one game’s revenue.
Redistribution requires a license copy, notices of modifications and preservation of applicable attribution and NOTICE material. Bundling weights with a game and downloading them later create different distribution workflows, but neither removes commercial-use conditions. Check the covered entity and terms; contact Liquid AI about a commercial license if the threshold is exceeded.
LFM Open License v1.0 text and commercial conditions
What a small integration trial should measure
A first prototype could choose the strategy of one NPC from text state. Start an experiment at 1–2 Hz, with additional requests on scene changes. This is a small test design, not a measured recommended operating frequency. Compare the same scene with the game alone, AI alone and both running together.
Record median and p95 request-to-action latency, cold-start delay, slow frames, peak VRAM and RAM, and CPU load. Separate short text, long text and image inputs, and check decisions affected by quantization. High output probabilities do not establish enjoyable NPC behavior; inspect oscillation and implausible repeated choices too.
Running on the same PC is plausible. The adoption criterion is whether the target PC preserves smooth rendering and delivers useful decisions on time, rather than one small file size or minimum latency. Testing that with large downloads and workload stress is a separate task from preparing this article.
Sources and scope
Publication and verification date: October 9, 2026; release announcement: October 7. Latencies are developer-reported, the memory table is arithmetic from nominal parameter counts, and figures and NPC designs are LATENT examples. Benchmark tables in the blog and evolving model cards differ, so their differing averages are not combined into one comparison.