The biggest speedup in local inference is one you never turned on
To work through your prompt, attention compares every token with every other token. A thousand tokens means a thousand-by-thousand grid of comparisons. Flash attention just refuses to build the grid.
Bigger means slower: more weights to drag out of memory for every token. Want more capability? Pay for it in speed. By now it feels like physics. But there is a hack!