Apple researcher, co-author of LLM in a Flash
Machine Learning Engineer · current
Keivan co-authored LLM in a Flash, the paper that showed how to run models bigger than a device's memory by streaming weights from flash. It is exactly the kind of systems insight that makes truly private, on-device AI possible for everyone.
We built this from your public work because we think it deserves celebrating. You did not ask us to, so the only fair thing is that you decide what happens to it. Claim it and it is yours to edit. Ask us to change something and we will. Ask us to take it down and it is gone within 72 hours — free, no account needed, and nobody will try to talk you out of it.