inference-server - 技术专题

相关标签
macosinference-servermlxapple-siliconopenai-apillmrustmachine-learningtransformersapi-server

Here are 129 public repositories matching this topic...

Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.

  • Updated Jul 28, 2026
  • Go