Shard - getting to 10× KV cache compression
A drop-in HuggingFace Cache for Llama-3.1-8B that makes KV memory about 10× smaller at 8K context, using different compression paths for K, V, and decode tokens.
Notes and essays on hardware/software co-design, FPGA systems, AI infrastructure, and practical engineering tradeoffs.
A drop-in HuggingFace Cache for Llama-3.1-8B that makes KV memory about 10× smaller at 8K context, using different compression paths for K, V, and decode tokens.