kahn1
Okura66/kahn1
High-throughput, sub-20ms System 1 decision engine on LLM logits. Zero text generation, calibrated probabilities, vLLM prefix caching.
SDK 与集成Python
- 星标
- 0
- 派生
- 0
审查来源
查看引用来源主题
jevdecision-enginefast-inferencellm-inferencepaged-attentionprobability-calibrationpythonqwen2-5