jev-bench
Running-Dolphins/jev-bench
Measure accuracy and calibration of Jev (TypeSafe AI's decision model) on public datasets: 12 business-like tasks, 7 experiments, one Python file.
Research & evaluationPython
- Stars
- 0
- Forks
- 0
Review source
View cited sourceTopics
jevbenchmarkcalibrationllmpython