laya-web: Laya running entirely in the browser as an 8-bit ONNX model
Live at laya-web.pages.dev. 524 MB of quantized English weights cached in the tab, onnxruntime-web over WebAssembly, and a parity page checked against the PyTorch reference.
Verified within 0.02 of the PyTorch reference.
# laya-web [Laya](https://huggingface.co/convaiinnovations/laya) — a System One decision model — running entirely in the browser, as an 8-bit quantized copy of the English checkpoint. **Live at [laya-web.pages.dev](https://laya-web.pages.dev)** · weights on [Hugging Face](https://huggingface.co/nvkudva/laya-web-q8) Laya is not a chat model. It reads a **state** — any text or JSON — scores the **options you enumerate**, and returns one calibrated probability distribution per question in a single forward pass. No sampling, no tokens out. ## Features - **Three question types.** `noul` (true/false), `choice` (pick one named option), `score` (expectation over ordered levels). - **Calibrated probabilities**, not raw softmax: temperatures fitted per question type and option count, plus an entropy-based confidence on every answer. - **Nothing leaves the tab.** onnxruntime-web over WebAssembly — no API key, no backend, and no request carrying your state anywhere. - **524 MB of weights, cached across reloads** in the Cache API. The page shows how much is cached and lets you delete it. - **Verified against the PyTorch reference.** `/parity.html` re-runs both gates in the browser: byte-identical token ids, then end-to-end probabilities within 0.02. - **Five worked presets** — email triage, moderation, routing, review scoring, escalation — each one state answered by three questions at once. ## How it works Every question becomes one sequence, with a `[MASK]` marker in front of each option: ``` [CLS] <type> question: <instructions> [SEP] [MASK] opt0 [MASK] opt1 … [SEP] <state> [SEP] ``` 1. **Tokenize** with transformers.js, straight from `tokenizer.json` — no model needed. 2. **Encode** the sequence with the quantized ModernBERT-large encoder (512 tokens). 3. *