すべてのプロジェクト

jev-ragcheck

gabazureus/jev-ragcheck

RAG evaluation with typed decisions: sentence-level hallucination verdicts with offsets and citations, passage and answer relevance, in one Jev call per answer. Benchmarked against Ragas, DeepEval, an LLM judge and HHEM.

研究と評価Python
スター
0
フォーク
0

掲載理由

掲載理由は準備中です

参照元を見る

トピック

jevbenchmarkfaithfulnesshallucination-detectionllm-as-a-judgellm-evaluationportuguesepython