すべてのプロジェクト

jev-bench

Running-Dolphins/jev-bench

Measure accuracy and calibration of Jev (TypeSafe AI's decision model) on public datasets: 12 business-like tasks, 7 experiments, one Python file.

研究と評価Python
スター
0
フォーク
0

審査時の参照元

参照元を見る

トピック

jevbenchmarkcalibrationllmpython