workflowreadBluesky: @simonwillison.netAug 06, 2026
Performance data that could inform model selection or serving config.
workflowreadBluesky: @simonwillison.netAug 05, 2026
Cloud/closed-model news - context for the local-first value proposition.
hardwarewatchllama.cpp CommitsAug 06, 2026
Performance-sensitive backend path - could affect local throughput.
hardwarewatchllama.cpp CommitsAug 06, 2026
llama.cpp commit - read only if it touches your serving path.
hardwarewatchllama.cpp CommitsAug 05, 2026
Performance-sensitive backend path - could affect local throughput.
Field Note
Jun 05, 2026
A current field note on why the 3x3090 serving path moved through pipeline parallelism, not tensor parallelism, for the tested 32B AWQ model.
Research
Jun 05, 2026
A public-safe read on Gemma 4 12B as a local sensory preprocessor: useful for seeing, hearing, and structuring observations without turning into an action system.
Hardware
Jun 05, 2026
A measured note on 300W bursty inference, lower caps for sustained runs, and why power-cap sweet spots are workload-specific.
Hardware
Jun 05, 2026
A field-report read on 14x RTX 3090 agent serving, EXL3, FP8 KV cache, Aphrodite, and why concurrency is becoming the local hardware metric.