Java port of Laya with choice/score/noul, a script-detecting router that picks the checkpoint, HuggingFace Hub fetching and bit-level parity with the Python ONNX output.
GitHubTools & apps
115 builds · page 3 of 3
Java port of Laya with choice/score/noul, a script-detecting router that picks the checkpoint, HuggingFace Hub fetching and bit-level parity with the Python ONNX output.
GitHubTools & apps
Tokenizer, mmBERT-base encoder and decision head in one binary for the 322M multilingual checkpoint. No Python, torch or BLAS; forward pass checked against a reference.
GitHubTools & apps
Loads original Laya safetensors with Candle and runs choice/score/noul with fused kernels, tiled attention and GEMM dispatch tuned for Apple Silicon. Optional f16 on Metal.
GitHubTools & apps
No Python, no torch, no ONNX export. Loads the ungated safetensors directly and runs any of the three checkpoints on CPU, Metal or CUDA.
GitHubTools & apps
ggml-based C++ runtime for all three checkpoints. Native tokenization, decision heads and JSON output; optimized FP16/BF16 CUDA on RTX with FP32 accumulation, plus Vulkan.
GitHubTools & apps
npm package that runs Laya with ONNX Runtime, no Python or PyTorch at runtime. Same request shape as TypeSafe's system_one API; output matches Python to four decimals.
GitHubTools & apps
Validated Core ML ports of all Laya checkpoints. One short multilingual decision in 4.98 ms P50 on M3 Max with ANE FP16, at 2.78x better energy per decision than MLX FP16.
GitHubTools & apps
X postTools & apps
X postTools & apps
Asking everyone to fine tune https://t.co/mwPvWGPY2n Just use any kinda agentic harness like claude code, codex etc. Change the code and use free kaggle GPUs to train your own use cases. Base model has limitations
There’s something cute about giving a probabilistic classifier a freshman probability exam. So I gave @typesafeai Jev, SemIf, and Laya ~1000 adapted questions on probability and countability from Berkeley exams Jev: 83.7% SemIf: 61.6% Laya: 31.2% A little stress test for the
X postDocuments & data
I benchmarked Jev 1.13.0 vs open-weights Laya on 751 identical questions. Jev sweeps triage/guardrails/moderation, goes 1.000 on 5-language intent. Laya wins agnews + mnli at $0. Full data:
X postLLM routing & guardrails
Tested Jev vs Laya (local, MLX) on 40 support tickets, EN + Thai. Jev 90–100%. Laya zero-shot 30–80%. Laya's faster, but not close on accuracy yet.
X postMultilingual
Laya shipped right after Jev claiming to be faster for AI agent decisions. curious if it was accurate too I built a POC same Google ADK triage agent, same 30 labeled tickets, only the engine swapped. Jev: 422ms, 53% accuracy. Laya: 152ms 10% accuracy Faster not more accurate
X postSupport & triage
Spent a morning testing a small open-weights local model (Laya-MLX, fully offline on a MacBook Air, $0 per call) on real tasks: inbox triage, tagging 100 SAP decks, categorizing blog posts, and a prefilter pattern. Honest numbers in this thread.
X postSupport & triage
Benchmarked @mizorewww’s Laya-MLX against @FeatherlessAI Simple Jev on my M4 Pro inside @OpenClaw: Laya: 85% @ 27.8ms Jev/Qwen3-4B: 90% @ 1,167.6ms The stack is rules → Laya → Jev → @ollama 27B → cloud. 14/20 decisions resolved before Jev. @typesafeai was right about System
X postLLM routing & guardrails
I've added Laya and https://t.co/flwFG40TX4 (@michael_chomsky) support to leanest, Jev is good but having other options to try out seems nice. Also new og image (I hate og image caching so bad) https://t.co/OgVzZg8CVE

X postLLM routing & guardrails
I ported laya to my specific hardware (6900XT) with the help of my hermes agent. Check it out here
X postTools & apps
Wow, Laya (https://t.co/PhmZAwTZeX) inference in Jolt (Clojure running on Chez Scheme)
X postTools & apps
Built a hypermedia browser on Laya-MLX — a small model that, like Jev, doesn’t generate text; it just returns a probability for each option you give it. The next link is chosen from HAL’s _links and what ALPS declares those links mean. No routes defined up front. cc/ @mamund
X postLLM routing & guardrails
I'm publishing this workflow! a 3-model fully local pipeline (prompt enhancement + Laya decision-making) that generates an image in ~37s (fast af) Should run great on any RTX 3080+ https://t.co/i11GcPJbIm

X postLLM routing & guardrails
Built laya-drift to visualise and monitor agentic-drift today. It continuously measures how far an OpenCode session has drifted from its original plan using Laya + Jensen-Shannon divergence https://t.co/eB9DNlQksp

X postLLM routing & guardrails
I made a postgres plugin for the Laya decision models, using @Kelwing's Rust framework for running them
X postDocuments & data
I forked Needle (the Chrome extension) to use Laya instead of Jev. Fully local AI browser search. Try it out here: https://t.co/SYB7HQfya9 It is not a semantic search in strict terms, it will match your search to semantically similar results.
X postDocuments & data
Update: I trained Laya non-step for close to 24 hours. No where close to jev 0-shot :|
X postTools & apps
جالب شد بازی اسنیک. ورودی لوکیشن میوه و هد و بوردر هارو میگیره و با مدل اوپن سورس #Laya به صورت لوکال که شبیه #Jev هست دستور میگیره که لوکیشن بعدی کجا بره بالا پایین چپ راست. تجربه جالبی بود اینم لینک پروژه تو گیت هاب ممنون میشم با 🌟 حمایت بکنین: https://t.co/dcQcs6WljY

X postTools & apps
The engineer and the trainee. Adaptive MPC knows the physics. Laya, a small decision model, learned from examples. Alone, Laya lands 0/90. With MPC vetoing about one in three of its moves, it gets 90/90, through a 60% engine loss. Fly it in your browser:
X postTools & apps
A vision variant of Laya on Hugging Face: the same choice/score/noul question shape, with an image as the state.
GitHubModeration & safety
ONNX export of the Laya checkpoint so it runs anywhere onnxruntime does, including browsers and edge boxes. An fp16 variant followed a day later.
GitHubTools & apps
GGUF conversions of both the English and multilingual Laya checkpoints for llama.cpp-style local runtimes.
GitHubMultilingual
A browser-use agent that picks the next action with Laya, trained and evaluated on the Mind2Web web-task benchmark. Weights on Hugging Face.
GitHubLLM routing & guardrails
MLX port of Laya with performance work. Under 1 GB of memory on-device; the demo plays Snake at 60 decisions per second on an M3 Max.
Nandakishor’s dev.to write-up on building decision models a year before Jev, RLCD training against proper scoring rules, and why Laya is open.
ResourceTools & apps
Gradio demo of Laya on Hugging Face: email triage, phishing, guardrails, ticket routing, RAG filtering and multilingual routing, no install.
SiteSupport & triage
The model itself. Typed choice/score/noul questions over any state in one forward pass, 100+ languages, Apache 2.0 weights.