madewithlaya

Catalog / Tools & apps

0051GitHub

laya.cpp: standalone C++ inference with CUDA and Vulkan backends

ggml-based C++ runtime for all three checkpoints. Native tokenization, decision heads and JSON output; optimized FP16/BF16 CUDA on RTX with FP32 accumulation, plus Vulkan.

Throughput measured on an RTX PRO 6000 Blackwell at batches 1–8; precision checked against the Python baseline.

Open source ↗ github.comcostfree, localtime-
lkarlslund/laya.cppREADME ↗
# laya.cpp

Standalone C++ inference for Laya typed decisions using ggml, with optimized
CUDA execution on NVIDIA RTX GPUs and a Vulkan FP32 backend. Model loading, Unicode/BPE tokenization, transformer
inference, decision heads, and JSON output all run natively.

All three checkpoints are supported:

| Variant | Context | Encoder width / layers | Tokenizer |
|---|---:|---:|---|
| `english` | 512 | 1024 / 28 | NFC byte-level BPE |
| `multilingual` | 1024 | 768 / 22 | Metaspace BPE with byte fallback |
| `typed-decisions` | 1024 | 1024 / 28 | NFC byte-level BPE |

Strict FP32 is the default. The optimized CUDA path uses exact checkpoint FP16
weights, paired activation components, FP32 accumulation, and fused packing
kernels. Enable it with `--tensor-core-fp32 --flash-fp32`. Native mixed BF16 is
available with `--bf16` using the validated CUDA 13.0 / cuBLAS 13.1.0 build profile.
Correctness is checked against the baseline at matching precision; see
[precision comparisons](docs/precision.md).

## Performance

Measured on an **NVIDIA RTX PRO 6000 Blackwell (96 GB), capped at 450 W**, using
all 250 fixed acceptance questions at batches 1, 2, 4 and 8. Throughput is
**questions/second**; higher is better. Python denotes the original implementation.
The 16-bit mode is **BF16**.

| Model | Batch | Python BF16 | laya.cpp BF16 | Python FP32 | laya.cpp FP32 |
|---|---:|---:|---:|---:|---:|
| english | 1 | 149.2 | 365.8 | 147.9 | 341.7 |
| english | 2 | 267.5 | 585.6 | 201.6 | 420.6 |
| english | 4 | 459.8 | 761.1 | 232.9 | 437.2 |
| english | 8 | 662.5 | 809.8 | 231.7 | 385.7 |
| multilingual | 1 | 178.9 | 485.4 | 191.7 | 475.3 |
| multilingual | 2 | 320.7 | 816.0 | 299.8 | 610.0 |
| multilingual | 4 | 550.2 | 1,164.0 | 388.5 | 673.4 |
| multilingual | 8 | 828.1 | 1,246.4 | 408.0 |

Also filed under Tools & apps

  1. 0005

    laya-mlx plays Snake at 60 decisions per second

    介绍比Jev快50倍,在你设备上跑的laya-mlx! 只在你的设备上占用最高1G内存 Laya是一个开源的类似于Jev的,基于文本输出概率的分类系统 我将其移植到MLX,并且做了一些性能优化! 视频中就是这个模型在我的本地M3Max上玩贪吃蛇 这个模型能够以每秒决策60次的速度玩贪吃蛇! https://t.co/31KGUiNunb

    @mizorewww · Tools & apps · free, local · 60 decisions/s

  2. 0112

    Laya on Axera AX650 and AX8850 edge NPUs, under 70 ms

    AXERA-TECH's conversion of all three Laya checkpoints for its edge AI chips, reported at latency under 70 ms per decision on AX650/AX8850.

    AXERA-TECH · Tools & apps · free, on-device · < 70 ms on AX650 / AX8850

  3. 0099

    laya-mcp (PerryLink): a warm sidecar that fixes Laya's silent truncation and constant noul

    MCP server plus HTTP sidecar: token-budget preflight that reports what would be cut, a persisted calibration store, structured errors, and a noul fix (P(true) 0.5 → 1.0).

    @PerryLink · Tools & apps · free, local

  4. 0097

    Gomoku against Laya-MLX: you play black, the model picks white's move from six candidates

    Local 15x15 Gomoku on Apple Silicon. Rules in Python, board in the browser, every AI move one real Laya choice over up to six candidates, probabilities shown. No Gomoku training.

    smile-magic · Tools & apps · free, local