jev-bench
Running-Dolphins/jev-bench
Measure accuracy and calibration of Jev (TypeSafe AI's decision model) on public datasets: 12 business-like tasks, 7 experiments, one Python file.
研究与评测Python
- 星标
- 0
- 派生
- 0
审查来源
查看引用来源主题
jevbenchmarkcalibrationllmpython