Creating Intention, Consuming Implementation
Intention-Implementation and the Mechanical Oracle
The role of the mechanical oracle was swiftly introduced at the end of the previously chapter.
Dim the lights in proportion to the quality of the oracle, not the model. A more capable model has the potential to introduce less errors but does not distinguish them as such. Being able to verify cheaply is perhaps the winning ticket to giving the factory a shot at winning.
What defines the quality of an oracle can be as such:
Soundness
- Does a pass actually mean correct implementation, or it's not obviously wrong? Is the tax calculator tested for all known rules, or are we just checking that it never divides by zero?
Coverage
- Self-explanatory, an alternate view is how does it constrain the intention? If the deviation between intention and implementation is obviously here, then the coverage is not good enough.
Cost and latency
- A good oracle can't lag: the faster a loop closes, the better the feedback cycle. The cost here is not just the operational costs of the agents, but the cost of shipping too.
Independence from generator (model)
- If the implementation agent has already misunderstood the intention, then it can't be checking its own work. Tests can be written by a different model with an independent context; tests are best verified by that set of independent eyes.
There is also a scale to which the team can aid the oracle to get effective:
Proofs, or a typed programming language/system
Property-based testing (easiest level after clearing basic compilation issues)
Plain, mechanical refactoring when the opportunity comes - changing implementations without test regression is a cheap way of checking.
Input from the people in the team - a senior teammate's taste is a guiding board for both junior teammates and the oracle to learn from, before the latter gains more independence.
Other than the scale, there is also one more feature that can help - literally testing in production as long as it's cheap and error-free for reversions. Having no push-back from real users of the system is a solid indicator for the oracle, that 'no news is good news'.
Remember - frontier models move at a much faster pace that the team can practically evaluate their oracles, which are built specific to intent. The rational allocation of effort is favorable towards oracles, because the other axis improves for free. Writing the property test is higher leverage than writing the better prompt.
Creating Intention
This is one way of describing how agentic-friendly teams work now: the creation lies not in code but in guidelines, prompts, specifications (GPS - I wonder if I can work this into another article...).
Combining the two points, creating intention means to use the store of knowledge as the source for thinking what needs to be built. Having a lower cost of creation suggests that the intentions can be done easily, such as a standardized specification format where pattern than novelty is easier replicated.
Consuming Implementation
This is also one way of describing how said teams can check themselves: having themselves - with mechanical oracles - to code review in order to understand whether the feedback loop holds.
Combining the two points, consuming implementation means to build up the store of knowledge with the processing of the materialized impact. Having a lower cost of consumption suggests that the implementations should tend to be simple to reason and understand.
Impedance: Signal or Noise
Just to link it back to the earlier idea of impedance - a self-described source of friction and losses when the intention-implementation translation happens. It's easy to frame these as noise, but it's also equally worthwhile to treat the corrections as signals.
The delta between the intention and implementation is a gap that can be closed, and that closure is something to train the mechanical oracle for. These corrections are more than just that - taste is also up for triage and identification. As mentioned earlier, a change in frameworks might make past conventions blunt, even if they don't immediately become deprecated (and irrelevant).
Harvesting these corrections as a corpus is also a pointer to potential oddities in the code repository - either they are accepted (and becomes a rule to enforce for instead of against), or the repository is treated.
AI disclaimer: Have bounced ideas off, so while text is still my own, the way I arrived at some of them may contain minor influences from such interactions.