kahn1
Okura66/kahn1
High-throughput, sub-20ms System 1 decision engine on LLM logits. Zero text generation, calibrated probabilities, vLLM prefix caching.
SDKs & integrationsPython
- Stars
- 0
- Forks
- 0
Review source
View cited sourceTopics
jevdecision-enginefast-inferencellm-inferencepaged-attentionprobability-calibrationpythonqwen2-5