kahn1
Okura66/kahn1
High-throughput, sub-20ms System 1 decision engine on LLM logits. Zero text generation, calibrated probabilities, vLLM prefix caching.
SDK と連携Python
- スター
- 0
- フォーク
- 0
審査時の参照元
参照元を見るトピック
jevdecision-enginefast-inferencellm-inferencepaged-attentionprobability-calibrationpythonqwen2-5