jev-ragcheck
gabazureus/jev-ragcheck
RAG evaluation with typed decisions: sentence-level hallucination verdicts with offsets and citations, passage and answer relevance, in one Jev call per answer. Benchmarked against Ragas, DeepEval, an LLM judge and HHEM.
研究と評価Python
- スター
- 0
- フォーク
- 0
掲載理由
掲載理由は準備中です
参照元を見るトピック
jevbenchmarkfaithfulnesshallucination-detectionllm-as-a-judgellm-evaluationportuguesepython