jev-agent-failure-benchmark
TokenTrim/jev-agent-failure-benchmark
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).
Research & evaluationPython
- Stars
- 2
- Forks
- 0
Topics
jevpython
TokenTrim/jev-agent-failure-benchmark
Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).