jev-test
clduab11/jev-test
Pre-registered benchmark: can a 2B local model (Gemma 4 E2B) answer web questions without making things up when a decision model (TypeSafe Jev) makes every call? SearXNG for search, MemPalace for verbatim memory, seven arms including open local judges. Spec and thresholds fixed before any run.
研究と評価Python
- スター
- 1
- フォーク
- 0
審査時の参照元
参照元を見るトピック
jevbenchmarkgemmahallucinationmempalaceragsearxngsmall-language-models