All projects

jev-agent-failure-benchmark

TokenTrim/jev-agent-failure-benchmark

Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).

Research & evaluationPython
Stars
2
Forks
0

Topics

jevpython