High-performance LLM inference in C++ enabling local AI on CPUs and Apple Silicon. The foundational engine powering most local AI tools.
llama.cpp is currently fading in Local AI — signals suggest reduced community momentum; check the GitHub commit cadence before committing to this tool.