Model-agnostic NPU+GPU+CPU inference engine for AMD Strix Halo. FastFlowLM reverse-engineered — 19 architectures, 46+ 1BP models, 5 backends. GGUF/ONNX/Q4NX/1BP. Zero Python. MIT.
vulkan mit-license quantization mamba inference-engine model-agnostic cplusplus-23 ai-inference local-llm one-binary gguf open-source-ai amd-strix-halo npu-inference fused-engine xdna-2 zero-python amd-native 1bp-format ternary-inference
-
Updated
Aug 10, 2026 - C++