Shoehorn — Quantize LLMs to your hardware
notactuallytreyanastasio.github.io
CLICK TO OPEN Interactive browser-based tool for quantizing large language models to fit your personal machine. Scans Hugging Face's most-downloaded models, estimates which ones fit your specific hardware (Mac, GPU, custom configs), visualizes memory budgets, then provides a local CLI app and web interface to run the chosen model via llama.cpp. Combines practical utility with a clean, hands-on approach to a niche problem.