Shard ↗︎
KV-cache compression for Llama-3.1-8B using RoPE-aware PCA and quantization. 10.0× smaller at 8K and 11.2× at 32K, with the decode tradeoffs measured rather than hidden.
Projects across AI inference, FPGA systems, silicon tooling, and computer architecture.
KV-cache compression for Llama-3.1-8B using RoPE-aware PCA and quantization. 10.0× smaller at 8K and 11.2× at 32K, with the decode tradeoffs measured rather than hidden.
FPGA timing-analysis copilot that parses Vivado timing reports, groups failing paths, and classifies issues across logic depth, routing delay, fanout, pipelining, and constraints. Analyzed 128 failing endpoints with –7.614 ns worst negative slack.