laya-rust: pure-Rust Laya inference on candle, CPU/Metal/CUDA
No Python, no torch, no ONNX export. Loads the ungated safetensors directly and runs any of the three checkpoints on CPU, Metal or CUDA.
# laya
Rust inference for [Laya](https://huggingface.co/convaiinnovations/laya), a non-autoregressive
typed-decision model of the kind TypeSafe calls a
[System One model](https://typesafe.ai/blog/introducing-system-one-models-and-jev). You give
it a **state** (text or JSON) and a set of **typed questions**; it returns typed answers with
calibrated probabilities in a single forward pass. It never
generates text, so there is nothing to parse and nothing to hallucinate.
Pure Rust on [candle](https://github.com/huggingface/candle) — no Python, no torch, no ONNX
export step. Runs on CPU, Metal, or CUDA.
## Getting the weights
The root checkpoint is Apache-2.0 and ungated, so no token is needed:
```sh
mkdir -p models/laya-base/encoder models/laya-base/tokenizer
B=https://huggingface.co/convaiinnovations/laya/resolve/main
for f in model.safetensors encoder/config.json tokenizer/tokenizer.json \
tokenizer/tokenizer_config.json rl_agent_config.json; do
curl -sL -o models/laya-base/$f "$B/$f"
done
```
That is ~847 MB. The repo also carries two sibling variants under
[`multilingual/`](https://huggingface.co/convaiinnovations/laya-multilingual) and
[`typed-decisions/`](https://huggingface.co/convaiinnovations/laya-typed-decisions); this crate
loads any of them, since the layout is identical, so point `--model` at the directory you
downloaded. `typed-decisions/` is the fine-tuned one, marked
`"fine_tuned": true` with temperatures near 1.0 rather than near 2.0, and it is the sensible
default for typed decisions.
## CLI
```sh
cargo build --release
./target/release/laya \
--state-file examples/ticket.txt \
--questions examples/questions.json
```
```json
{
"model": "rl-agent",
"answers": {
"department": {
"type": "choice",
"choice": "billing",