All projects

jev-bench

Running-Dolphins/jev-bench

Measure accuracy and calibration of Jev (TypeSafe AI's decision model) on public datasets: 12 business-like tasks, 7 experiments, one Python file.

Research & evaluationPython
Stars
0
Forks
0

Review source

View cited source

Topics

jevbenchmarkcalibrationllmpython