typed-decision-bench
4nt0ineB/typed-decision-bench
Bench of typed decision models: Jev vs OpenJev vs Laya, small local LLMs and cheap hosted LLMs on the same zero-shot classification tasks, in English and French.
Research & evaluationPython
- Stars
- 0
- Forks
- 0
Why it is listed
This small benchmark compares Jev with OpenJev, Laya, and local/hosted models on shared zero-shot English/French classification tasks; the authors limit it to one dataset and question type, and the metrics were not reproduced.
Topics
jevbenchmarkcalibrationenglishfrenchlayallm-evaluationmassive-dataset