すべてのプロジェクト

foreman-jev-evaluation

MahdiHedhli/foreman-jev-evaluation

JEV evaluation (FM-JEV-01): replay-only, advisory-only evaluation of Foreman-style supervision with TypeSafe Jev. Measures how much safety comes from the model versus a deterministic evidence gate. No worker authority.

研究と評価Python
スター
0
フォーク
0

審査時の参照元

参照元を見る

トピック

jevagent-supervisionai-safetyevaluation-harnessllm-evaluationtypesafepython