Vol. I · No. 123THU, AUG 20, 2026
Archive

The Archive

Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.

ExLlamaV3 Major Updates!

ExLlamaV3 adds Gemma 4 support, improved caching, and DFlash optimization for faster LLM inference on consumer hardware.

··

Markdown browser for LLMs

TextWeb: markdown renderer for LLM web browsing via MCP, reducing vision model overhead with native text reasoning.

··

Session limit frustration!

User reports frustration with Claude API session and weekly rate limits consuming quota rapidly during continuation work.

··

First MCPs, then Skills, now Memories are next

This was a really good talk, especially for anyone who's built things like the Karpathy wiki, Serena, or SQLite databases as memory for Claude. For any senior devs out there, are you spotting the solutions already implemented in distributed systems being reused? If many agents are working in parallel, how do you get them from stepping on each others toes? I can imagine logical clocks, consensus, deduplication, idempotency, and eventual vs causal consistency being applied. If you're on the Anthropic team, I'm curious how much different distributed systems algos were experimented with.

··
30 stories