h.jk

Glass to Silicon - Lens Kit

We started this series with a discourse about focal-length and resolution in photography, and how that might translate onto agentic characteristics and performance.

Let's see how we can adopt the approaches to lenses collection (a lens kit!) to the team's choice of agents.

Bag of Primes

A prime lens has a fixed focal length and often features a larger maximum aperture. Some might opt for a smaller maximum so that they can be physically much smaller - a surprising equivalent to smaller models with smaller context windows in the past. That is if you extend the depth-of-field perspective that the aperture controls with how much light is gathered at a given snapshot, comparable to the scope-per-read.

Outside of the technical characteristics, prime lenses are also favored for the artistic constraints that they impose on the photographer - they have to 'zoom with their feet', and one can become very experienced with the focal length over time. Even back in the film days, it is possible to develop an intimacy with a focal length that one can take photos without peering through the viewfinder.

If one prefers to pack more focal lengths within reach, then they will have to carry more lenses. This is the 'bag of primes' approach. Three lenses are the typical starting point - a wide, standard, and telephoto. Those without analysis paralysis may scale up to five or more - creativity is the limit!

There are parallels with the choice of agents in these regards. Dedicating a shared experience of a specific agent and thinking effort allows the team to know where it excels and where it falls off. Getting different people acquainted with different agents reduce the analysis paralysis, and that flexibility can be a positive factor.

However, a bag of primes still necessitates the changing of the lens at hand - attempt to use the 'telephoto' agent when you are doing a wider security audit will miss important vectors at breadth, but at least you know something is off. This changing of the lens is often paid in operational overheads - commands to swap, runbooks to confirm, rebuilding contexts when agents can't pool.

The plurality of agents might also lead to wanton accumulation of prompts and outdated technical guides when agents don't share or even disagree with approaches. Cross-cutting tasks, such as moving from a coding agent to running QA, doesn't fit a single prime lens. A team running this configuration will have to carefully balance the sharpness of a pool of task-focused agents with the bag (collective harness) they have to carry.

Superzoom

As it names implies, this is meant to replace a bag of primes with one dependable workhorse. One trade-off is they typically have smaller variable apertures across the zoom range, even those with a constant aperture doesn't feature the widest apertures that prime lenses get.

Modern frontier models, a close analogue to superzooms, typically feature class-leading context windows though. Mid-task changes flow more reasonably so, without requiring the team to amass a larger operational runbook for lens-swapping.

The second trade-off is the perceived image quality. Modern lens design has done great improvements, but when it comes down to it, superzooms are described as 'just as sharp as prime lenses', not the other way round. Again, modern frontier models tend to scale better, but subtle concurrency bugs or deep security reviews might become blind spots if they are optimized for breadth not depth in their training.

Zooming into a misjudged focal length also brings about confidently wrong answers - an inadequate prime lens is simply out-of-reach, but the superzoom might just assume an attack vector isn't possible because it wasn't asked to look there. The focus (pun intended) here is more relevant instead - coupled with the smaller apertures, superzooms require more light through shutter speed or sensitivity (ISO) adjustments to become confidently correct with prime lens.

A superzoom might seem like a good fit for heterogeneous codebases (like a monorepo) but only just - one frontier agent model on one repository is the ultimate simplification but there is a small, calculated technical risk of running thinner contexts when the model has to adjust for different languages and frameworks.

Why Not Both

Practically, this is the final accommodation a learned photographer settles for. A generalist superzoom that does either:

And a pair of primes (thereabouts) for:

Here, we make the bridge to the other half of this series, the silicon of 'Glass to Silicon'. We will examine comparable characteristics of image sensors and dynamic ranges with the agentic usage.

AI disclaimer: Have bounced ideas off, so while text is still my own, the way I arrived at some of them may contain minor influences from such interactions.

#glass-to-silicon