I stopped asking one model to play every role. Pokémon Red showed me how to cast a crew by benchmark and route each job with vLLM.

Brian Douglas
Head of DX at Continue
Oakland, CA
I work at the Paper Compute Company.
Previously founded Open Sauced (joined Linux Foundation 2024) and led Developer Advocacy at GitHub.
Host of "Open Source Ready" and "The Secret Sauce" podcasts. Passionate about mentoring new open-source contributors.
Latest Posts
My OpenClaw chief-of-staff agent was burning $80 a day in Anthropic API costs. The fix was one config line, and the only reason I found it was observability.
Skills are a great way to teach an agent a pattern. I got curious whether the agent could just learn the pattern, so I tried training a model while also figuring out what RL actually is.
I built a tool that fans out parallel Claude Code agents to fix lint errors across a codebase. It taught me that the hard problems aren't prompts or models. They're isolation, observability, and memory.
How a DeepMind paper turned my Pokemon agent from 'watch and tweak' into 'run and measure.'