Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.
-
Updated
Oct 3, 2026 - Swift
Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.
Quality-first local MiniMax H3 video-series generation and loopback API for dual RTX 4090 workstations, with native audio, references, P8/P9 continuity, and preserved artifacts.
Research preview for reproducible MoE routing and memory oversubscription on consumer GPUs. Not a production inference engine.
Flash-backed mixture-of-experts inference on Apple devices with Swift/MLX.
Deploying personal-grade production models on a consumer-grade PC
DualDeadline adds separate gate/up and down-projection transfer deadlines to exact MoE offloading, with H200 validation and Triton-optimized predictors.
Stream PyTorch model blocks from NVMe or pinned CPU memory with a bounded GPU residency budget. LoRA-finetune models far larger than host RAM on one GPU - Qwen3-VL-32B and gpt-oss-120b train in under 6 GB of VRAM.
Research artifact for Memory-Sovereign Inference: Output-Exact Execution Beyond Full Residency
A research toolbox for running and studying LLMs on constrained hardware, with memory, placement, inference, and benchmarking methods.
To associate your repository with the model-offloading topic, visit your repo's landing page and select "manage topics."