kevala: Laya and Kev in any web page, zero-dependency Rust to WebAssembly with WebGPU
One import, no server. Dependency-free Rust compiled to Wasm with WebGPU kernels; pulls a pinned int8 pack from Hugging Face and keeps it in the browser. Live Tetris included.
kevala</h1>
<p align="center"><b>Ask questions about text and get answers from small models running on the user's own GPU.</b></p>
<p align="center">
<a href="https://bvolpato.github.io/kevala/#/tetris"><img src="docs/tetris.gif" alt="Tetris played live by the Laya decision model on WebGPU: the model scores every landing spot and presses the keys to get there" width="760"></a>
</p>
<p align="center">
<a href="https://bvolpato.github.io/kevala/">Live site</a> ·
<a href="https://bvolpato.github.io/kevala/#/playground">Playground</a> ·
<a href="https://bvolpato.github.io/kevala/#/tetris">Tetris</a> ·
<a href="https://www.npmjs.com/package/kevala">npm</a> ·
<a href="examples/">Examples</a> ·
<a href="docs/architecture.md">How it works</a>
</p>
kevala runs System 1 decision models, [Laya](https://huggingface.co/convaiinnovations/laya) and
[Kev](https://github.com/jaredpalmer/kev) today, inside the browser. You give it a piece of text or
JSON and typed questions (`noul` for yes/no, `choice`, `score`), and it returns a probability for
every option of every question in one forward pass. There is no model server, no API key, and no
data leaving the tab.
The engine is Rust with zero dependencies, compiled to WebAssembly, plus WebGPU kernels written for
these models. The browser runtime is a few plain ES modules. The first load downloads a pinned int8
pack of the model from [Hugging Face](https://huggingface.co/bvolpato/kevala-packs) and keeps it in
the browser, so there is nothing to host. If the pack is unreachable, kevala converts the original
checkpoint in the browser instead.
## Quick start
From a CDN, in any page:
```html
<script type="module">
import { Kevala } from "https://cdn.jsdelivr.net/npm/kevala@latest/js/src/index.js";
const kevala = await