LLMの応答を3〜10倍高速化!KVキャッシュを賢く再利用する省エネエンジン — LMCache

LMCache/LMCachePython8.7k
GitHubで見る →

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

技術情報

言語

Python

ライセンス

Apache-2.0

最終更新

2026-06-13

スター数

8,742

フォーク数

1,298

Issue数

311

トピック

amdcudafastinferencekv-cachellmpytorchrocmspeedvllm

過去のトレンド履歴

関連リポジトリ