jev-no-enem
patryckalves/jev-no-enem
Reproducible benchmark evaluating TypeSafe AI's Jev (System One paradigm) on Brazil's ENEM 2025 standardized exam. Evaluates typed decision-making, domain-specific accuracy, and RLCD uncertainty calibration against open LLM baselines with an interactive GitHub Pages dashboard.
研究与评测Python
- 星标
- 0
- 派生
- 0
审查来源
查看引用来源主题
jevbenchmarkbrasilenemllmrlcdsystem-one-modelspython