jev-test
clduab11/jev-test
Pre-registered benchmark: can a 2B local model (Gemma 4 E2B) answer web questions without making things up when a decision model (TypeSafe Jev) makes every call? SearXNG for search, MemPalace for verbatim memory, seven arms including open local judges. Spec and thresholds fixed before any run.
研究与评测Python
- 星标
- 1
- 派生
- 0
审查来源
查看引用来源主题
jevbenchmarkgemmahallucinationmempalaceragsearxngsmall-language-models