jev-agent-failure-benchmark
TokenTrim/jev-agent-failure-benchmark
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
研究と評価Python
- スター
- 2
- フォーク
- 0
トピック
jevpython
TokenTrim/jev-agent-failure-benchmark
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).