Contextomics
A call to study the efficient production of economically valuable software.
When ChatGPT rolled out its Gmail integration back in the heady days of GPT-4 I imagined it as the start of an amazing personal assistant. I sent in my first query, asking it to identify marketing emails in my inbox and ranking them by the most prolific sender so we could set up smart filters. The query took 40 minutes, an insane number of tool calls, and an impressive Python script, all to produce a disappointing 70% answer. Worst of all it then threw away the Python script upon the start of a new conversation. It brought to mind Elon Musk’s likening of legacy rockets to building a 747 flying across the ocean, then throwing the plane away.
Claude Code / Codex are so far ahead of GPT-4 to be alien tech, but they still produce the same feeling in me watching them work. While it’s impressive that they can figure out the things they do by trial and discovery, watching them reinvent the wheel for the nth time gets old fast. It’s intuitive that asking Claude to review and spot type violations itself is a waste of tokens, just let it tool call tsc/ty and respond to the output, then it can spend its review cycles on higher value issues. Instead of asking it to look at a backend API and write the client interface, use an OpenAPI schema + generator it can invoke with a single tool call, 5k tokens → 45 tokens. An astute reader might recognize these are all practices that made human coders more efficient too! The difference is that there was always a shortage of data / professional study around exactly which practices were truly better for human coders. Even when data was collected we were fundamentally comparing groups of individuals, anyone could always dismiss results as “skill issues” as to why their favorite tool they spent 10 years mastering performed poorly in a study, tribalism and cargo culting against reason and logic, ironic for the software discipline.
We now have a historic opportunity in front of us to correct this. With coding agents we can set up real experiments and engage in a new study of Contextomics 1 . It’s now cheap to ask questions, do ORMs hurt or hinder development velocity, which architecture patterns really matter, plain HTML/JS vs React, does typing matter for Python code, should we write all software going forward in Rust? We don’t just want to measure these practices statically , e.g. tokens needed to build a todo app, good economically profitable software is dynamic , once the initial app is built how fast can it be refactored/extended to chase ever evolving market needs? I want to see software benchmarks move away from pure measures of correctness, and towards measures of efficiency across time. Implement feature A, then how fast can you refactor to enable a new use case while preserving the original feature?
Contextomics matters most for our libraries, languages, and tools. We need to be considering that the primary user going forward will be agents. The question to ask is what can be done to make APIs token efficient? Anecdotally I’ve seen that Claude prefers flat over nested APIs, whether they are HTTP/CLI, I’m excited to experimentally confirm/disprove this. Onwards towards a future of better software!