MLX core contributor, quantization and fast kernels
MLX engineer · current
Alex's quantization work, from quantized attention to 3 and 6 bit weights, is why big models fit on ordinary Macs at all. He does hard numerical engineering in public pull requests where anyone can learn from it.
We built this from your public work because we think it deserves celebrating. You did not ask us to, so the only fair thing is that you decide what happens to it. Claim it and it is yours to edit. Ask us to change something and we will. Ask us to take it down and it is gone within 72 hours — free, no account needed, and nobody will try to talk you out of it.