All projects

kahn1

Okura66/kahn1

High-throughput, sub-20ms System 1 decision engine on LLM logits. Zero text generation, calibrated probabilities, vLLM prefix caching.

SDKs & integrationsPython
Stars
0
Forks
0

Review source

View cited source

Topics

jevdecision-enginefast-inferencellm-inferencepaged-attentionprobability-calibrationpythonqwen2-5