laya-candle: Rust inference with fused Metal kernels tuned for M1 Pro
Loads original Laya safetensors with Candle and runs choice/score/noul with fused kernels, tiled attention and GEMM dispatch tuned for Apple Silicon. Optional f16 on Metal.
# laya-candle Native Rust inference for [Laya](https://github.com/NandhaKishorM/laya) using [Candle](https://github.com/huggingface/candle). Pass text or JSON and typed questions to receive choices, ordinal scores, or boolean probabilities. Loads original Laya safetensors directly, with no Python runtime. ## Usage Use Rust 1.95 (pinned in `rust-toolchain.toml`) and a native C/C++ build toolchain. Clone the source and run with CPU float32, the default backend: ```sh git clone https://github.com/b0xtch/laya-candle.git cd laya-candle cargo run --release -- predict \ --model convaiinnovations/laya \ --state-file examples/state.json --questions examples/questions.json ``` On Apple Silicon: ```sh cargo run --release --features metal -- predict \ --model convaiinnovations/laya --device metal \ --state 'Please refund the duplicate charge.' \ --questions examples/questions.json ``` Use `--dtype f16` for reduced precision on Metal. The final scorer and calibration remain float32, but reduced precision can change probabilities and close decisions. Metal uses fused kernels, tiled attention, and GEMM dispatch tuned for M1 Pro; other Metal devices use Candle's matrix dispatch. CPU requires float32. `--features accelerate` enables Apple's CPU BLAS backend. CUDA (`--features cuda`, `--device cuda`) and bfloat16 are available but have not been validated on hardware for this project. Install the CLI locally with `cargo install --path . --locked` (add `--features metal` for Apple Silicon), then invoke `laya-candle predict ...`. The library and CLI are experimental; the 0.1 API may change. The first Hub load downloads the selected checkpoint. Subsequent loads reuse the Hugging Face cache. `--model /path/to/checkpoint` loads local weights; `--offline` uses cached files wi