ToolRadarHQ

Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac

TurboFieldfare is a Swift and Metal inference engine built specifically to run 4-bit quantized Gemma 4 26B on M-series Macs — at a memory footprint that would normally buy you a much smaller model. The 2 GB RAM figure is the headline, and if it holds up under scrutiny, it meaningfully lowers the bar for serious on-device inference on consumer hardware. What sets it apart from generic llama.cpp wrappers is the native stack: Swift plus Metal means it bypasses the Python runtime entirely and speaks directly to Apple's GPU. That is a real architectural choice, not a cosmetic one. The honest reservation is that this is early-stage, single-developer work — the repo will need community stress-testing before you stake a production pipeline on it. But as a local inference experiment for a weekend, the RAM claim alone is worth verifying. -> Best for: AI engineer or indie hacker who wants to run capable models locally on Apple silicon without buying more RAM
More like this