~1,000 Berkeley probability exam questions: Jev 83.7%, SemIf 61.6%, Laya 31.2%
Abhijay gave three probabilistic classifiers adapted probability-and-countability questions from Berkeley exams as a stress test.
Abhijay gave three probabilistic classifiers adapted probability-and-countability questions from Berkeley exams as a stress test.
Seven zero-shot extraction and classification families demoed, benchmarked and cross-compared, with sizes, licences and costs in one table.
@umstek · Documents & data · free for the local rows
@umstek
GitHub
Documents & data
free for the local rows
-
Frozen protocol, raw predictions, calibration and latency for Jev 1.13, PrismNLI-0.4B and Laya. Laya ties Jev on dair-ai/emotion and trails by 10–33 points elsewhere.
@elcronos · Documents & data
@elcronos
GitHub
Documents & data
-
-
Six classifiers over SST-2, AG News and more, reporting accuracy, macro-F1, ECE, p50/p95 latency, throughput and cost side by side. Laya runs locally on MPS or CPU.
@dhruvmehra · Documents & data · $0 for the local rows
@dhruvmehra
GitHub
Documents & data
$0 for the local rows
-
Anthony Alcaraz's talk demo routes typed decisions through Jev and shows Laya's answers beside them. On the TSMC run Laya rejected every true claim; that's the point.
@AnthonyAlcaraz · Documents & data · free (Laya shadow); Bedrock for reasoning
@AnthonyAlcaraz
GitHub
Documents & data
free (Laya shadow); Bedrock for reasoning
-
There’s something cute about giving a probabilistic classifier a freshman probability exam. So I gave @typesafeai Jev, SemIf, and Laya ~1000 adapted questions on probability and countability from Berkeley exams Jev: 83.7% SemIf: 61.6% Laya: 31.2% A little stress test for the Show more