madewithlaya

Catalog / Tools & apps

0049GitHub

Laya-CoreML: 4.98 ms decisions on the Neural Engine, 2.78x better energy than MLX

Validated Core ML ports of all Laya checkpoints. One short multilingual decision in 4.98 ms P50 on M3 Max with ANE FP16, at 2.78x better energy per decision than MLX FP16.

Open source ↗ github.comcostfree, localtime4.98 ms P50 / 5.31 ms P95 on M3 Max
mizorewww/laya-coremlREADME ↗
# Laya-CoreML

**Open-weight typed decisions on Apple Silicon. Core ML, Neural Engine, zero generated tokens.**

[PyPI](https://pypi.org/project/laya-coreml/) · [Hugging Face weights](https://huggingface.co/aac6fef/laya-multilingual-coreml-ane) · [中文](https://github.com/mizorewww/laya-coreml/blob/main/README.zh-CN.md)

A real Laya model plays Snake locally, with visible probabilities, score, length,
latency and safety interventions. The GIF replays a recorded Core ML run at **1× speed**.
The game uses explicit planner features and a visible cycle safety layer.

The complete active Snake loop sustained **49.1–50.0 decisions/s** across three
uncapped 600-step episodes, with zero deaths and two safety interventions.
[Game-loop timings and paced-rate limits](https://github.com/mizorewww/laya-coreml/blob/main/docs/SNAKE_BENCHMARKS.md)
include rendering serialization; terminal painting is excluded.

**One short multilingual decision: 4.98 ms P50 / 5.31 ms P95 on M3 Max with ANE FP16.**
The same experiment measured **2.78× better whole-system energy per decision** than
compiled MLX FP16. A separately validated W8 palette variant reached 4.88 ms and
3.19× energy improvement. These are single-question results, not full Snake frame
times; the requested 10× improvement was not achieved.

## Run the demo

Apple Silicon · macOS 15+ · Python 3.11–3.13.

```bash
pip install 'laya-coreml[demo]'
hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/snake
laya-coreml-snake --model ./models/snake
```

Download once, then play offline. No PyTorch, Transformers or MLX is needed for
inference. The terminal needs 104 columns × 35 rows. Space pauses; ↑/↓ changes
speed; R resets; Q quits. First-time Core ML initialization can take tens of seconds.

[Controls, recording and video e

Also filed under Tools & apps

  1. 0005

    laya-mlx plays Snake at 60 decisions per second

    介绍比Jev快50倍,在你设备上跑的laya-mlx! 只在你的设备上占用最高1G内存 Laya是一个开源的类似于Jev的,基于文本输出概率的分类系统 我将其移植到MLX,并且做了一些性能优化! 视频中就是这个模型在我的本地M3Max上玩贪吃蛇 这个模型能够以每秒决策60次的速度玩贪吃蛇! https://t.co/31KGUiNunb

    @mizorewww · Tools & apps · free, local · 60 decisions/s

  2. 0112

    Laya on Axera AX650 and AX8850 edge NPUs, under 70 ms

    AXERA-TECH's conversion of all three Laya checkpoints for its edge AI chips, reported at latency under 70 ms per decision on AX650/AX8850.

    AXERA-TECH · Tools & apps · free, on-device · < 70 ms on AX650 / AX8850

  3. 0099

    laya-mcp (PerryLink): a warm sidecar that fixes Laya's silent truncation and constant noul

    MCP server plus HTTP sidecar: token-budget preflight that reports what would be cut, a persisted calibration store, structured errors, and a noul fix (P(true) 0.5 → 1.0).

    @PerryLink · Tools & apps · free, local

  4. 0097

    Gomoku against Laya-MLX: you play black, the model picks white's move from six candidates

    Local 15x15 Gomoku on Apple Silicon. Rules in Python, board in the browser, every AI move one real Laya choice over up to six candidates, probabilities shown. No Gomoku training.

    smile-magic · Tools & apps · free, local