全部项目

jev-no-enem

patryckalves/jev-no-enem

Reproducible benchmark evaluating TypeSafe AI's Jev (System One paradigm) on Brazil's ENEM 2025 standardized exam. Evaluates typed decision-making, domain-specific accuracy, and RLCD uncertainty calibration against open LLM baselines with an interactive GitHub Pages dashboard.

研究与评测Python
星标
0
派生
0

审查来源

查看引用来源

主题

jevbenchmarkbrasilenemllmrlcdsystem-one-modelspython