AI news, with the context that matters.

Technology & Research · ·

Can Liquid AI d1 run on the same PC as your game?

d1-3B and the experimental 600M return decisions in one pass. What the published 8 ms really measures, and how to budget memory and integrate decisions without stalling a game.

An enemy NPC needs to choose whether to pursue, hide or call for help. That decision need not involve generating a long answer. Liquid AI released d1-3B and the experimental d1-omni-600M on October 7, 2026: models that return typed decisions for questions supplied by an application.

Integration paths exist. Running smoothly on the same PC as a game is a separate question, and the published 8 ms on an RTX 4090 does not settle it. This article examines public information as of October 9, 2026. LATENT has not run the models or measured them alongside a game.

Liquid AI release announcement

Key takeaways

  1. d1 returns choices or scores in one forward pass without generating prose. It can be evaluated for NPC action selection or intent classification; the game still executes movement and attacks.
  2. The published RTX 4090 result of 8 ms uses a warm, optimized short-question workload. It does not establish latency for long inputs or competing rendering workloads; headroom on the target PC needs separate measurement.
  3. Official GGUF releases provide a llama.cpp route and smaller quantized files. File size is not runtime memory usage, and the weights use lfm1.0 with commercial conditions.

A model for NPC decisions

d1-3B accepts text and images. The early research 600M release accepts text with images or text with audio, but not images and audio together. Its audio training uses requests between an English speaker and an assistant, with clips cut at 30 seconds. Generalization to Japanese in-game speech needs evaluation.

Question type Example game design
choice Choose among allowed actions such as retreat, wait and call for help.
noul Return a probability that an utterance asks for help, without composing a reply.
score Rate an ordered scale such as calm, alert and critical.

3B input and output specification

600M input and audio limitations

The shared idea with our Strands Decider article is to return a selection without text generation and leave execution to the application. Here, the interesting extension is to images and, for 600M, audio. The backbones also differ: d1-3B builds on LFM2.5-VL-3B, while 600M uses a bidirectional encoder. Comparing an older Decider timing with these results on different hardware and inputs would not establish a winner.

Related: how Strands Decider works

The model selects; the game acts

An illustrative NPC design, not a tested integration

1 Select state

Extract health, distance and the recent utterance needed for the decision.

2 Score options

Send allowed actions to d1. Provide movement coordinates and dialogue separately.

3 Validate rules

Check that the action remains valid, then pass it to the behavior tree or state machine.

Figure 1. A high probability does not make an action legal or safe. The game excludes unavailable actions.

What 8 ms measures and why input size matters

These are developer measurements, not concurrent gaming tests. The RTX 4090 row uses BF16 and the median of 20 warm runs. Single questions use CUDA graphs through model.compile; the reported uncompiled result is 16 ms. A new input shape incurs initial kernel selection or compilation work.

RTX 4090 workload Published latency
Short single question 8 ms
Three questions over one state 21 ms
3.4K-token state 102 ms
384 px image 17 ms

Published results and conditions: d1-3B model card

One pass still has to process its input. At 102 ms, the 3.4K-token workload takes about 12.8 times the 8 ms result. Likewise, throughput of 475 states per second with 64 packed states does not mean an isolated request returns in about 2 ms. Batch throughput and response latency differ. No speed results accompany this 600M release, so smaller size alone does not establish that it is faster.

A frame at 60 fps lasts about 16.7 ms, but rendering and game logic already need part of that time. Subtracting 8 ms from 16.7 ms does not establish a concurrent budget. Inference and rendering share GPU compute, memory bandwidth, power limits and queues. Average fps can stay acceptable while some individual frames become much slower.

Separate weight size from runtime budgets

Start with a weights-only estimate. Using the reported 3.12 billion and 587 million parameters, multiply parameter count by bytes per value. The table is neither measured VRAM usage nor a recommended RAM capacity. GB means one billion bytes; GiB means 2 to the power of 30 bytes.

Estimated full weights 16bit 32bit
d1-3B 6.24 GB / 5.81 GiB 12.48 GB / 11.62 GiB
d1-omni-600M 1.174 GB / 1.09 GiB 2.348 GB / 2.19 GiB

Separately, the official GGUF releases list a 1.67 GB 3B Q4_K_M main file and a 407 MB 600M Q8_0 main file. These are distribution sizes, not all-modality totals or peak runtime usage. Media projectors, input processing, intermediate computations and runtime allocations add to them. Quantization format and CPU/GPU placement also matter.

Official 3B GGUF and distribution sizes

Official 600M GGUF and distribution sizes

Distribution combination Sum of published file sizes
3B Q4_K_M main file plus Q8 image projector 1.67 GB + 583 MB ≈ 2.25 GB
600M Q8_0 main file plus Q8 media projector 407 MB + 263 MB ≈ 670 MB
600M F16 main file plus F16 media projector 764 MB + 441 MB ≈ 1.205 GB

These totals add the displayed distribution sizes. They exclude runtime placement and working allocations, so they are not minimum free-RAM or free-VRAM specifications.

3B file listing

600M file listing

Budget three resources on one PC

Fitting in memory and meeting a deadline are separate conditions

GPU memory

Game textures and render buffers plus model weights and temporary allocations.

GPU time

Rendering plus inference. Asynchronous work still shares compute and bandwidth.

CPU and RAM

Physics, audio and input processing plus CPU inference or model loading.

Figure 2. The model file alone cannot determine how much VRAM is sufficient alongside a game.

Suppose a game uses 6 GiB of an 8 GiB GPU, leaving 2 GiB. The estimated 5.81 GiB of full 16-bit 3B weights would not fit there. But a 1.67 GB quantized file being smaller than 2 GiB still does not prove it will fit at runtime. Moving work to the CPU can reduce GPU demand, while introducing RAM, CPU-time and transfer constraints. This is a budgeting example, not a measurement of a particular game or PC.

Keep inference off the rendering path

A sensible first target is a strategy change every few seconds or a decision at a dialogue boundary, rather than per-frame aiming or collision detection. Submit a request when health falls, an enemy is lost or the player finishes speaking. There is no need to turn an already deterministic distance check into a language-model request.

In an illustrative design, rendering continues with the current behavior while a separate inference process receives a state snapshot. Tag the request with NPC ID, state version and deadline; discard stale results. Keep the existing behavior tree running if a reply is late. Limiting each NPC to one pending request helps prevent a queue of decisions about obsolete situations.

Frequency matters. Twenty NPCs making two requests each per second produce 40 requests per second. Even at a hypothetical serial 8 ms per request, that adds up to 320 ms of processing per second. This is not measured GPU utilization under contention. If requests take 100 ms, the same serial worker cannot keep up. A design can reserve frequent decisions for important visible NPCs and use conventional logic for distant ones.

Multiple questions can share a state, but avoid mixing knowledge that different NPCs should not share. When the game can supply a compact internal state, its input and processing are easier to control than capturing and interpreting the screen again. For voice input, measure the player-visible delay including recording and preprocessing.

Python and llama.cpp integration paths

Official Python examples use Transformers and PyTorch. The requirements are Transformers 5.14 or later for 3B and 5.15 or later for 600M. Examples use float32 on CPU, bfloat16 for 3B on GPU and float16 for 600M on GPU. Do not apply the same dtype indiscriminately: the 600M card warns against bfloat16.

Loading requires trust_remote_code=True, which executes supplied model code. Review and pin the code and revision before adopting it. This investigation did not install dependencies or download weights.

The models are not Python-only. The official GGUF READMEs show loading them into llama.cpp llama-server and sending state and questions to /v1/systemone. A 600M question must fit in one batch; the example sets -b 4096 -ub 4096. Images and audio also require the appropriate projectors.

A local HTTP connection from Unity or Unreal is a possible integration design, not evidence of an official engine plugin. llama.cpp merged 3B support in 88dcc46 on October 7 and 600M support in a657f7e on October 8. Use a build containing the relevant change and verify the target OS, dedicated endpoint and required modalities. The minimum packaged release number and Windows runtime behavior have not been tested here.

llama.cpp 3B support PR

llama.cpp 600M support PR

Do not infer that ordinary Ollama or LM Studio chat interfaces expose this decision API just from generic Hugging Face usage suggestions. Check the dedicated call documented in the author-maintained README, rather than a text-generation chat interface.

Calling 3B through llama.cpp

Calling 600M through llama.cpp

Commercial conditions for shipping a game

The model cards specify lfm1.0, the LFM Open License v1.0. Although based on Apache 2.0, it is not Apache or MIT. Commercial use has a US$10 million annual-revenue threshold; commercial use by a legal entity exceeding the threshold is not licensed under this agreement. Do not assess eligibility using only one game’s revenue.

Redistribution requires a license copy, notices of modifications and preservation of applicable attribution and NOTICE material. Bundling weights with a game and downloading them later create different distribution workflows, but neither removes commercial-use conditions. Check the covered entity and terms; contact Liquid AI about a commercial license if the threshold is exceeded.

LFM Open License v1.0 text and commercial conditions

What a small integration trial should measure

A first prototype could choose the strategy of one NPC from text state. Start an experiment at 1–2 Hz, with additional requests on scene changes. This is a small test design, not a measured recommended operating frequency. Compare the same scene with the game alone, AI alone and both running together.

Record median and p95 request-to-action latency, cold-start delay, slow frames, peak VRAM and RAM, and CPU load. Separate short text, long text and image inputs, and check decisions affected by quantization. High output probabilities do not establish enjoyable NPC behavior; inspect oscillation and implausible repeated choices too.

Running on the same PC is plausible. The adoption criterion is whether the target PC preserves smooth rendering and delivers useful decisions on time, rather than one small file size or minimum latency. Testing that with large downloads and workload stress is a separate task from preparing this article.

Sources and scope

Publication and verification date: October 9, 2026; release announcement: October 7. Latencies are developer-reported, the memory table is arithmetic from nominal parameter counts, and figures and NPC designs are LATENT examples. Benchmark tables in the blog and evolving model cards differ, so their differing averages are not combined into one comparison.

Primary source: d1-3B model card

Primary source: d1-omni-600M model card