flash-attention - 技术专题

相关标签
natural-language-processingchinesepretrained-modelslarge-language-modelsllmflash-attentioncudacuda-kernelscuda-democuda-toolkit

Here are 179 public repositories matching this topic...

From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

  • Updated Jun 5, 2026
  • Python