Logo
The Agent Engineer
Home
Articles
About
Search
Subscribe
Logo
Search

Inference Speed

Deepdive into LLM Inference Speed & Optimization techniques

The biggest speedup in local inference is one you never turned on

Jul 1, 2026

•

2 min read

Optimization techniques

+2

The biggest speedup in local inference is one you never turned on

To work through your prompt, attention compares every token with every other token. A thousand tokens means a thousand-by-thousand grid of comparisons. Flash attention just refuses to build the grid.

Yatharth Lakhera
Yatharth Lakhera
Speculative decoding isn't free. It's a bet you can lose.

Jun 28, 2026

•

2 min read

Optimization techniques

+2

Speculative decoding isn't free. It's a bet you can lose.

Best case, several tokens for the price of one. Worst case, slower than doing nothing. And what you're wagering on is how predictable your output is.

Yatharth Lakhera
Yatharth Lakhera
Your bandwidth is fixed. Your speed isn't.

Jun 24, 2026

•

2 min read

Optimization techniques

+2

Your bandwidth is fixed. Your speed isn't.

Pop-up wardrobes, tape-to-stream labs, and story prompts powering this week’s experiments.

Yatharth Lakhera
Yatharth Lakhera
How models get smarter without getting slower

Jun 22, 2026

•

2 min read

Optimization techniques

+2

How models get smarter without getting slower

Bigger means slower: more weights to drag out of memory for every token. Want more capability? Pay for it in speed. By now it feels like physics. But there is a hack!

Yatharth Lakhera
Yatharth Lakhera

The Agent Engineer

Deep dives on building AI agents that actually ship — Claude Code internals, production LLM patterns, and the stuff that breaks between a clean demo and real users.

© 2026 The Agent Engineer.
beehiivPowered by beehiiv