Columns · · Team LATENT
How practical is a local LLM you can actually rely on?
What can today’s Mac, DGX Spark, and Windows PCs offer someone who wants to own their AI? A practical look at Japanese quality, published performance evidence, VRAM and RAM, Strata, and what ¥1 million or ¥2 million buys.
There is a reason to want AI on your own computer. Drafts and work documents can stay with you instead of being sent out for every question. A model you like can remain available after a service changes its lineup. Owning AI means more than reducing a monthly bill.
The satisfaction of seeing a local model produce a reply wears off quickly. Can it write Japanese without repeated correction? Can it handle a long document within a tolerable wait? When asked to change code, can it carry the task through to verification? For daily use, these differences matter more than the price of the machine.
As of October 7, 2026, moving a bounded set of tasks onto your own hardware is realistic. Replacing everything you ask a frontier cloud service to do, with the same speed and effort, is a much less straightforward purchase. The ¥1 million and ¥2 million budgets below help expose that difference; they are not spending targets.
Key takeaways
- For Japanese drafts and bounded code changes, smaller models and the 27B class are useful candidates. A 16GB GPU constrains quantization and context choices; 32GB provides more room to keep a 27B-class model on the GPU.
- Mac and Spark offer large shared memory pools; a Windows GPU workstation offers a different balance. Strata expands the options for large MoE models, but its speed depends on memory bandwidth, quantization, and GPU caching as well as RAM capacity.
- Read published benchmarks together with their model, quantization, and input conditions. Judge the initial wait and completion time as well as generation speed, while retaining cloud options for demanding reasoning and long autonomous tasks.
From a model that runs to a tool for work
The meaning of usable depends on the job. Standardizing the wording of an internal document puts a premium on preserving meaning and responding quickly. Discussing a design calls for noticing conflicts between requirements. Both tasks produce fluent text, but they do not demand the same capabilities.
Start with work you repeat each week, rather than the largest model a machine can load: extracting decisions from meeting notes, proposing tests from a short specification, or editing your own notes. When the input and expected result are clear, it is easier to check the answer and decide whether local execution is worthwhile.
An open-ended request with little context makes failures harder to diagnose. You may have to distinguish a model limitation from quantization, missing information, or a tool configuration problem. More expensive hardware does not remove that work.
Japanese, code, and long documents need different checks
Japanese quality involves more than grammatical sentences. Register, omitted subjects, terminology, and the preservation of qualifications all matter. A small model may be useful for a fixed formatting task while demanding more checking when it rewrites a long explanation without changing the meaning. Hardware speed alone cannot solve that problem.
For a starting comparison, consider Gemma 4 12B QAT for lighter tasks and Qwen3.8-27B for a broader workload. Gemma 4 also offers the 26B A4B MoE. A4B describes the active parameter scale during generation, not the storage requirement of a 4B model. Parameter counts alone do not establish which is better at Japanese.
01 / Dense and MoE
Weights you store and weights you use at each step
An FFN transforms information for each token. This sketch shows one layer, with the blocks used for this token highlighted.
Dense
Input for one token
After shared operations such as attention, enter this layer’s feed-forward network (FFN).
Use the full FFN
Each token passes through the same complete FFN.
Store this FFN’s weights and use them in the computation.
MoE
Input for one token
After shared operations such as attention, enter this layer’s feed-forward network (FFN).
A router selects experts
The selection can change for the next token.
Expert 1
Expert 2
Expert 3
Expert 4
Expert 5
Expert 6
Expert 7
Expert 8
Use the two green experts now. Keep the other experts’ weights available for later selections.
Reading “26B A4B”
Gemma 4 26B A4B has roughly 26 billion parameters in total and uses about 4 billion per token. Its stored weights do not shrink to those of a 4B model.
Placement depends on the implementation, including GPU and system RAM. Parameter count and architecture alone do not determine Japanese quality.
Google estimates roughly 6.7GB to load the 4-bit 12B model and 14.4GB for 26B A4B, including 20% loading overhead. Those estimates are not bounds for arbitrary context lengths or multiple users. On a 16GB GPU, the 12B choice leaves more room; the 26B A4B choice needs closer attention to actual memory use, including context.
Gemma 4 variants and memory estimates
Qwen3.8-27B emphasizes coding, professional work, and longer agent workflows. Its reported benchmark results depend on evaluation settings, context, and reasoning configuration. They should not be treated as guaranteed results from a compressed local setup.
Qwen3.8-27B model card and evaluation conditions
Quantization reduces storage by representing weights with fewer bits. Unsloth lists Qwen3.8-27B UD-Q4_K_M at 16.5GB and UD-Q6_K at 22GB. A 16.5GB file being close to a GPU’s advertised 16GB is not evidence of a comfortable fit. Unit conventions differ, and the runtime also needs context state such as the KV cache and working memory.
02 / Model size and quantization
The same 27B model can have different file sizes
Distinguish how many parameters a model has from how their values are stored.
Qwen3.8-27B
Choose the storage precision for the same model of roughly 27 billion parameters.
UD-Q4_K_M
16.5GB ≈ 15.4GiB
UD-Q6_K
22GB ≈ 20.5GiB
The bars show published file sizes. Q4 and Q6 identify quantization formats; the files are not simply 27B × 4 bits or 27B × 6 bits.
Changing precision changes storage needs and can affect outputs. Check the answers on your own tasks when choosing a setting.
03 / Inference memory budget
Loading the weights is only part of the memory budget
Budget within each memory pool: GPU VRAM or a unified pool such as a Mac’s shared memory.
① Loaded weights
The allocation made by the runtime can differ from the downloaded file size.
② KV cache
Stores context state. Longer contexts and concurrent requests generally need more space.
③ Working memory
Temporary buffers and other workspace depend on the runtime, execution method, and added features.
④ Other uses and headroom
A unified pool also serves the OS and applications. Display output and other GPU tasks can consume VRAM.
Fit ① + ② + ③ + ④ within the capacity of the same memory pool.
Check free space before extending context
Weights can load successfully while leaving too little space for the subsequent KV cache.
Consider 32GB VRAM for a daily 27B model
The article’s Q6-class, 32K-context configuration is a starting point to verify, not a guarantee for arbitrary extras or concurrency.
Keep the memory pools distinct
Adding system RAM does not expand VRAM. Moving from a 4TB to an 8TB SSD does not expand these working memory pools.
This gives 32GB VRAM a concrete role if a 27B-class model will be your daily tool. Q6-class weights with around 32K context provide a reasonable starting configuration to verify, not a guarantee after adding vision components, speculative decoding, or concurrent requests. With 16GB, first decide whether to accept a smaller model, stronger compression, or partial CPU offload.
Actual quantized files and sizes
Writing one function and investigating a repository to change several files are different jobs. The latter also requires software for search, editing, testing, and recovery from failures. Connecting a model API does not reproduce a cloud coding agent’s complete working environment. Small changes with existing tests make a clearer starting point.
For long documents, accepted input length and reliably used information are different limits. A supported context window does not eliminate omissions. Selecting relevant chapters or files, or retrieving passages with sources, can reduce prompt-processing time and make the answer easier to verify.
How Strata uses VRAM and system RAM
A Windows PC has two relevant capacities: GPU VRAM and system RAM. A model and its working state that fit on the GPU can be processed there. Offloading parts to RAM introduces CPU work or transfers over PCIe. Expanding system RAM from 64GB to 128GB does not turn a 16GB GPU into a 128GB GPU.
Strata is interesting because it designs that division of work around a large mixture-of-experts model. An MoE selects some experts for each input instead of applying every expert’s weights equally at every step. Strata caches frequently used experts on the GPU and keeps others in RAM for CPU work. That architecture and placement allow it to handle a model larger than GPU memory.
Upstream Strata environments and inference approach
The main target is Qwen3.8-Flash-Next, not arbitrary large models. Its 125B core is accompanied by 51B of per-layer embeddings and a 4B multi-token-prediction component. PLE provides embedding lookups; MTP supports predicting subsequent tokens. Neither the 125B label alone nor a simple sum of these counts determines required RAM.
Qwen3.8-Flash-Next architecture
Strata’s model guide lists a 54.8GB main component for IQ3_S, with about 50GB of experts, plus a separate roughly 28.8GB PLE file kept on SSD. A 64GB system can fit it with little room for other applications. Moving to 128GB provides room for less aggressive quantization and daily applications; UD-IQ4_XS involves roughly a 94GB download and 59.5GB of experts.
Keep two uses of SSD storage separate: sparse PLE lookups and frequent expert-weight reads caused by insufficient RAM. The latter slows inference. Moving from a 4TB to an 8TB SSD primarily increases storage for models and documents; it does not resolve a RAM shortage.
Strata quantization sizes and low-RAM behavior
04 / Where the model lives
Adding capacity does not make all memory equally fast
Distinguish where weights are stored from which processor computes.
Conventional GPU inference
Windows PC / NVIDIA GPU
GPU · VRAM
Model weights + KV cache + working memory
Process on the GPU when everything fits
CPU · RAM
Some configurations keep offloaded weights here and compute on the CPU
More system RAM is not more VRAM
Strata divides work between CPU and GPU
A specialized path for supported large MoE models
GPU · VRAM
Attention and other operations, plus a cache of frequently used experts
Cache hit rate also affects speed
CPU · RAM
Store expert weights and compute work outside the GPU cache
Both capacity and memory bandwidth matter
SSD · Look up required PLE entries
Sparse PLE reads differ from repeatedly reading expert weights from SSD because RAM is short. The latter slows inference.
CPU and GPU share memory capacity
Mac Studio / DGX Spark・OEM
Unified memory
Model weights, KV cache, OS, and applications use one memory pool.
Mac uses MLX or Metal; Spark uses CUDA. Their execution software differs.
The execution platform is another boundary. Upstream Strata’s usual installation targets are x86 Windows and Linux; GB10 support comes through a separate Arm64 fork. This is not the same installation path on a Mac. Mac users instead consider MLX or llama.cpp with Metal, so Strata speed reports cannot be transferred directly into a Mac purchase argument.
Operations are still evolving. Version 0.1.40.1, released October 6, fixed quoted tool calls being executed and requests hanging during engine restarts. Those are concrete reasons to check more than throughput, including whether displayed examples remain text rather than actions. For file-changing tasks, use a recoverable working area and explicit review points.
What published benchmarks tell us about waiting time
When reading a benchmark, separate the wait for the answer to start, the rate at which text appears, and the time to finish. Fast generation can still follow a long wait for input processing or hidden reasoning.
05 / Response timing
Starting quickly and finishing the answer are different
From submission to completion. Box lengths are not proportional to time.
Read the input
Model loading, queueing, and prompt processing. Longer inputs can increase this wait.
Generate reasoning
Additional time when reasoning is enabled. Internal tokens may appear before the answer intended for the reader begins.
Display the answer
Distinguish how quickly the text appears from how long it takes to finish.
Time to first visible answer
Submission → start of C
Includes the waits in A and B.
After the answer starts
How quickly text appears
The rate of visible answer text, separate from internal reasoning.
Time to completion
Submission → end of C
Judge the wait on answers representative of your work.
A published Strata RTX 5090 test illustrates the distinction. It used Linux, 64GB RAM, v0.1.29, and IQ2_XS on September 30, 2026. Reasoning was disabled, with synthetic code-explanation inputs and 256-token outputs. The following values are medians of three runs per condition.
| Input tokens | Decode rate | First API text delta (TTFT) | Time to 256-token completion |
|---|---|---|---|
| 4,096 | 179.4tok/s | 0.974 s | 2.392 s |
| 32,768 | 175.7tok/s | 5.950 s | 7.390 s |
| 128,000 | 165.0tok/s | 22.262 s | 23.850 s |
Decode rates differ relatively little, but TTFT exceeds 22 seconds on the longest input. Here TTFT runs from the HTTP request to the first nonempty text delta, rather than the first screen display. Model loading is excluded, and GPU expert-cache hit rates were 97.8–99.7%. These speeds cannot be assumed for the 128GB Windows configuration below or higher-precision quantization.
RTX 5090 test conditions and complete results
On GB10, an October 4 test used an ASUS GX10 and the Arm64 fork, with IQ3_XXS, MTP enabled, reasoning off, synthetic inputs, and a 128-token output cap. Each condition ran once. Inputs of roughly 1.6K–75K tokens produced 37–42 tok/s; the short 431-token input produced 15.3 tok/s.
The same test recorded total wall times of 16.1 seconds for 431 input tokens, 8.6 seconds for 6,190, and 60.0 seconds for 75,047, each with 128 output tokens. Shorter input was not always faster, making prompts representative of everyday work particularly relevant.
GB10 fork measurements, including the short-input result
The two reports differ in quantization, software, prompts, output lengths, and cache conditions. They are neither a controlled 5090-versus-Spark comparison nor an evaluation of answer quality. They establish a fast 5090 configuration and a working GB10 execution path with measured rates. Read tok/s with the understanding that token boundaries vary by model and text.
Before buying, try the intended model with your own documents and questions, and judge answer quality alongside waiting time. If a smaller model handles the work, there is no need to spend the full ¥1 million. More VRAM or shared memory becomes valuable when the work calls for larger models or longer documents.
What Mac, Spark, and Windows PCs are good at
Mac Studio’s clear advantage is a large memory pool shared by CPU and GPU. A 64GB or 256GB configuration is not split into GPU-only memory and separate system RAM. The OS and applications share that capacity too, so the entire installed amount is not a model-weight budget.
The M5 Mac Studio was announced August 25 and released September 22, while the 512GB configuration is scheduled for late October. The 64GB Max and 256GB Ultra examples are released products, but that does not mean immediate delivery: the retailer lists incoming-stock estimates of five to six and nine to ten weeks respectively.
The quoted 40-core GPU M5 Max has a stated memory bandwidth of 614GB/s; M5 Ultra is rated at 1.2TB/s. Bandwidth matters when repeatedly reading large models. Dividing these figures by Spark’s 273GB/s does not yield a Japanese or English generation-speed ratio: computation, quantization, libraries, and context differ.
Apple’s M5 Mac Studio announcement
A Mac is a natural choice for someone who already works on macOS and wants to explore models supported by MLX or Metal. The value of 256GB is room for models and higher-precision weights that are difficult to place in 32GB VRAM. CUDA-only implementations and the upstream Strata installation call for a different platform.
MLX LM and llama.cpp provide starting points for checking supported models and execution methods. Sufficient capacity does not establish software support for a newly released model.
DGX Spark and related OEM systems are best understood as compact CUDA machines with GB10 and 128GB shared memory. They use an Arm CPU and Linux-based DGX OS. They make sense as a separate inference host accessed from a laptop or as a CUDA development environment. An advertised maximum parameter count does not establish acceptable latency or concurrency.
NVIDIA’s October 2 announcement was not the initial launch of the Spark family. The new 64GB configuration is OEM-only, scheduled for October 23 from $4,999. This research did not establish Japanese pricing, so a currency conversion is not presented as a domestic sub-¥1-million purchase option.
Availability and OEM scope of the new 64GB configuration
A Windows PC provides a practical route to GPU-resident models with Strata as an additional option. GPUs and SSDs are easier to change later, and the machine can also serve gaming or creative work. The trade-offs include a large GPU’s power and heat, drivers, and software support. Buying faster GPU computation and buying room for larger models are separate decisions.
What ¥1 million and ¥2 million actually buy
These tax-inclusive Japanese configuration examples were checked on October 7, 2026. They are neither a lowest-price list nor a comparison at guaranteed equal answer quality. The baseline SSD is 4TB. An 8TB drive is an option for storing many large models and documents, not a prerequisite for fast inference.
| Example configuration | Tax-inclusive price and availability | Reason to choose | Remaining constraint |
|---|---|---|---|
| Windows/RTX 5080 16GB Core Ultra 7 270K Plus RAM 64GB・4TB SN850X |
¥829,280 Options plus shipping; delivery unconfirmed |
Keep smaller models on the GPU; share the PC with creative work or games | Little headroom for GPU-resident 27B models, depending on quantization and context |
| Mac Studio / M5 Max 18-core CPU, 40-core GPU 64GB unified memory, 4TB SSD |
¥923,800 Incoming stock estimated 5–6 weeks after order |
Combine daily Mac work with experiments in the 27B class and beyond | No CUDA; not the upstream Strata execution path |
| HP ZGX Nano / GB10 128GB unified memory, 4TB SSD |
¥1,097,800 Listed in stock during research |
Use a compact, separate CUDA inference machine | Exceeds ¥1 million; requires Arm support and an appropriate inference stack |
| Windows/RTX 5090 32GB Core Ultra 7 270K Plus RAM 128GB・4TB SSD |
¥1,873,280 Options plus shipping; delivery unconfirmed |
Switch between GPU-resident 27B models and Strata | VRAM remains 32GB; four RAM modules run at DDR5-3600 |
| Mac Studio / M5 Ultra 30-core CPU, 64-core GPU 256GB unified memory, 4TB SSD |
¥1,939,800 Incoming stock estimated 9–10 weeks after order |
Prioritize capacity for larger models or higher-precision weights | Capacity alone does not establish chat speed or answer quality |
For a new machine under ¥1 million, first decide whether you need a Mac or a Windows GPU workstation. A 64GB Max has a clear role for exploring several models including the 27B class on macOS. If you already own a GPU with at least 16GB, establish useful tasks before rushing to replace it with another card of the same capacity. A smaller model is not a reason to exhaust the budget.
The 128GB/4TB Spark examples sit just above ¥1 million. NVIDIA’s own DGX Spark was listed at ¥1,078,000, but the retailer marked it sold out with the next shipment expected in mid-October. A lower listed price does not mean immediate availability. Deciding whether you need a compact separate CUDA machine makes comparison with a Mac or tower clearer.
NVIDIA DGX Spark price and expected shipment
At up to ¥2 million, a Windows machine with 32GB VRAM and 128GB RAM can be compared with a Mac’s 256GB unified pool. For coding centered on a GPU-resident 27B model, with Strata used when appropriate, the 5090 path is attractive. If larger-model experiments and less aggressive quantization matter more, the Ultra path has a clearer purpose. Neither purchase guarantees frontier-cloud equivalence.
Pay particular attention to the Windows 128GB configuration. The listed four 32GB modules run at DDR5-3600, versus DDR5-5600 for two 32GB modules in the 64GB option. More capacity comes with lower CPU memory bandwidth. For Strata’s CPU expert work, capacity alone does not establish a superior configuration. If two 64GB modules are available, prioritize checking motherboard support and actual operating speed.
As another reference, ark lists an RTX 5090 system with 128GB DDR5-5600 and a 2TB SSD at ¥1,698,000. That is not a matching 4TB quote, and component availability matters. Treat it as a place to confirm a two-module 128GB layout, the 4TB total, and delivery timing rather than an automatic choice based on the lower listed price.
ark’s 128GB configuration example
The Dospara 5090 total adds ¥1,369,980 for the base machine, ¥308,000 for 128GB RAM, ¥192,000 for a 4TB SN850X, and ¥3,300 shipping. It is not a formal cart quote or reserved inventory. Cases and cooling also contribute to price differences, so the totals are not pure comparisons of LLM performance per yen.
The 5080 example adds ¥534,980 for the base PC, ¥99,000 for 64GB RAM, ¥192,000 for a 4TB SN850X, and ¥3,300 shipping. It includes the standard 1000W power supply and 240mm liquid cooler. This too is a total calculated from listed options, not a confirmed cart quote.
Deciding what your own AI should do
Define the local workload by its properties rather than a model name. Formatting drafts, extracting information from short documents, searching personal material, and making small testable code changes have inspectable inputs and outputs. Batch work with flexible deadlines is another natural local use.
Retain a strong cloud model for difficult design decisions, research in unfamiliar domains, and long autonomous workflows. Local preparation can reduce what needs to be sent out: organize documents or draft an approach locally, then escalate a limited question. If information cannot leave the machine, decide whether it can be separated from the question or whether more human judgment is needed.
The cost extends beyond hardware and electricity. Downloads, checks after updates, and troubleshooting consume your time. Dividing a machine’s price by a cloud subscription fee does not settle which is cheaper. Conversely, control over where data goes and the ability to retain a model version have value that a subscription comparison cannot fully capture.
If AI democratization means everyone owning the strongest possible AI, hardware and operational costs remain substantial. Yet the portion of useful work that can be retained independently of a provider is growing. Choosing the model, update schedule, supplied documents, and delegated authority is meaningful control.
A practical build starts with a small daily job and enough memory for that job. Expand after judging both the wait for a complete Japanese answer and the effort needed to correct it. The useful first step toward owning AI is finding a task you want it to keep doing, before spending the full budget on a large machine.
Research scope and sources
This research article draws on public documentation and Japanese retail listings checked on October 7, 2026, Japan time. LATENT did not benchmark the hardware, independently score Japanese quality, purchase a system, or download large models. The selection advice is editorial judgment based on specifications, published measurements, and software support. Prices, options, stock, and delivery estimates can change.
Model capabilities and formats come from their developers, file sizes from distributors, and speed figures from the original measurement reports. The reported measurements are used only within the conditions stated in the article. Retail prices can be checked through the links in the configuration table.