A customer-support team is told to reduce average call time. The number falls. Some agents become clearer and faster, but others transfer difficult callers, rush explanations, or close cases that return the next day. The dashboard celebrates improvement while customers repeat their stories to three people.
The team did not refuse to optimize. It optimized the available definition of success. Recursive self-improvement makes that problem sharper because a system that gets better at optimization can pursue an incomplete objective with growing skill.
Precision needs a worthy target. A tighter group is only an improvement relative to the chosen aim. Photo by Mark Stuckey on Unsplash
Objectives are compressed judgments
Every metric leaves something out. Revenue omits costs and distribution. Test scores omit parts of education. clicks omit satisfaction. Delivery speed omits safety. We use measurements because reality is too rich to carry whole into a decision, but the compression creates room between the goal and its proxy.
Economist Charles Goodhart’s original observation concerned monetary policy: a statistical regularity tends to break down when used for control. The broader family of ideas now called Goodhart’s law appears whenever pressure on a measure changes the behavior that once made the measure informative. People teach to the test, managers time transactions around reporting periods, and recommendation systems learn that outrage keeps a screen active.
Capability answers “How effectively can the system pursue an objective?” Direction asks whether the objective deserves pursuit.
The two can improve independently. A delivery network becomes faster and more reliable while its labor practices deteriorate. A political campaign becomes excellent at predicting which message will move a poll while making public argument less truthful. An engagement model learns individual preferences while gradually shaping those preferences toward whatever keeps engagement easiest.
Direction is not solved by adding more metrics. A balanced scorecard can reduce one blind spot and introduce five new incentives. Some goods resist clean measurement because they are qualitative, long-term, relational, or partly known only to the people living with the consequences. Judgment remains necessary.
Who gets to say?
“Better” always contains a point of view. Better for the customer may be worse for the worker. Better quarterly performance may weaken a company ten years later. Better national efficiency may reduce local freedom. Conflicts among real goods cannot be erased by technical vocabulary.
Watch the omitted person
When a metric improves, ask whose experience is absent from the data and who bears a cost outside the boundary of the system being optimized.
That doesn’t make measurement futile. It makes plural evidence and correction essential. Quantitative signals can reveal patterns intuition misses. Qualitative accounts can reveal harms a dashboard averaged away. Independent review can challenge the people who chose the target. Reversibility limits the damage when an objective turns out to be poorly specified.
Time changes the judgment too. A hospital intervention may raise short-term cost while preventing years of illness. Maintenance looks inefficient until the machine that was not maintained stops. Systems tuned to rapid feedback often discount slow goods, so the evaluation window is itself part of the objective.
A mature improvement process also distinguishes means from ends. Efficiency, scale, prediction, and control are powers. They can serve health, learning, friendship, craftsmanship, public order, or countless other goods, but they cannot choose among those goods on their own. A system that rewrites its method faster than people can reconsider its purpose may become impressive and less worth having.
The question “better at what?” should arrive before optimization and remain after every success. A target is never vindicated merely because the arrows begin landing closer to its center.