jev-bench
Running-Dolphins/jev-bench
Measure accuracy and calibration of Jev (TypeSafe AI's decision model) on public datasets: 12 business-like tasks, 7 experiments, one Python file.
研究と評価Python
- スター
- 0
- フォーク
- 0
審査時の参照元
参照元を見るトピック
jevbenchmarkcalibrationllmpython