全部项目

jev-agent-failure-benchmark

TokenTrim/jev-agent-failure-benchmark

Benchmarking Jev (Typesafe.ai) against a strong LLM on the Who&When Pro agent-failure-attribution benchmark (text subset).

研究与评测Python
星标
2
派生
0

主题

jevpython