KO한국어로 읽기
Economy경제사상
✓Fact-checked✓Code-verifiedvalidate.pyPublished
Note잔존 MAJOR 5(전부 '덜 논증됨'·본문이 미완을 스스로 노출 → 비차단) = ESG 범위 38% 경쟁가설 판별 부족 · 급소 정의가 착지에서 확장 · 완주사례↔신기여사례 상보 분포 · 축1 통념 스틸맨 부분 · 독자 자리 회수 부분. ★EN판 미동기 — 라이브 published/blog.en.md가 이번에 폐기한 44% 합산 라벨·'정도의 차이' 자기항복을 유지(사용자 결정 사항).

When a Bonus Rides on the Dashboard — A Metric Stops Reflecting the World and Becomes an Engine

2026.07.12·24 min read

Picture someone at the wheel. If a bonus rides on the speed number on the dashboard, the driver drives to raise the number rather than to actually go faster. He picks the downhill stretches; he floors the accelerator on roads where there is no hurry. Here the dashboard does not reflect speed. It produces it. The moment the thing being measured can read its own score and something rides on that score, measurement clocks out of its job as a camera and becomes an engine.

A thermometer does not change the temperature, and a desk does not stretch because you put a ruler to it. There are places where measuring plainly leaves its object untouched.

The person who set the camera against the engine is the sociologist Donald MacKenzie. Some seventy years ago an economist said that a good theory is not a photograph that copies reality but an "engine of analysis" that drives inquiry forward. MacKenzie turned that line on finance and put it on the cover of his book. An Engine, Not a Camera. A financial model, he argued, is not a camera that quietly photographs the market but an engine that, as it runs, pushes the market itself into a different shape.

That engine is what this piece takes apart. And once it is apart, the first question to ask in front of a metric changes. "Is this number accurate?" is not the first question.

Not a Belief but a Device — Reflexivity and Performativity

The claim that "measurement changes its object" is not itself new. The observer effect, where behavior shifts under watching, and the self-fulfilling prophecy, where a rumor makes itself true, are familiar to everyone. A bank run is the case: a rumor that "that bank is about to fail" pulls the deposits out, and the withdrawals manufacture the very insolvency they feared. We call this phenomenon, in which belief pushes reality through a feedback loop, reflexivity. Common sense stops here. Measurement does change behavior, yes—but that is a matter of belief or of misuse, and the act of measuring is itself neutral.

What we take on here lies one layer deeper. Think of the officiant at a wedding. "I now pronounce you married" does not describe a marriage that already exists. The words bring the marriage into being. In a 1955 lecture the philosopher of language J. L. Austin picked out the utterances that, instead of recording a fact, perform an act in the saying of it, and named them performatives.

When that idea crossed into economics it became performativity: formalized instruments and metrics lodge themselves in practice and constitute their object. Once a formula is lodged in the market as the tool that prices things and sheds risk, the market gets pulled toward the formula. Reflexivity and performativity are not strangers. Performativity is the larger umbrella, and reflexivity is one branch inside it. The landmark study demonstrating the performativity of measurement also names the self-fulfilling prophecy as one component of that feedback.

A skeptical reader can push here. "Isn't this the self-fulfilling prophecy with unfamiliar vocabulary swapped in?" There is one place where they part. Reflexivity requires belief—it is the structure by which a false belief makes itself true. Performativity works no matter what the participants believe, and even when the belief is correct. The key is not belief but the device. This piece does not push aside the lineage of the self-fulfilling prophecy. It only lays one more layer on top of it: the device.

Belief upstream of the device · Belief has not vanished. It has only moved upstream, into the designer's decision to lodge that device in the world—and the final question of this piece takes aim at exactly that upstream.

The distinction splits in practice. The KOSPI's plunge on the day of Samsung's earnings release was reflexivity, a forward-looking narrative outrunning and amplifying the hard numbers (covered in the memory-cycle piece). What we take on here is the larger pattern that contains that reflexivity: the phase in which a metric hardens into outright infrastructure and molds the world into its own shape. The Black-Scholes model we come to later is such a case. The market was dragged into that shape not because traders believed the formula but because they took it up as a working tool.

Raising an Object That Wasn't There — The Day an Index Invented "the Market"

The prototype of performativity sits in a number we hear every day: the stock index.

On 3 July 1884, Charles Dow ran a single figure in the financial newsletter he published—the average of the closing prices of eleven stocks. The number was 69.93. Twelve years later the Dow Jones Industrial Average, rebuilt out of twelve industrial stocks, opened at 40.94. It looks like idle arithmetic.

But until Dow made the choice to "take an average," no single object called "the market" existed. There was only a miscellaneous heap of scattered individual stocks. The average, as a metric, raised that heap into one object with a temperature and a water level. The sentence "the market rose today" holds only after Dow's arithmetic.

An objection comes straight back. "An index reflects the market; it does not make it. Dow transcribed what the market did, and index funds only track prices that active investors have already set. Prices exist independently, and the metric is just a thermometer." That's true of origins and of limits. But the engine lives in the feedback.

Once a metric becomes infrastructure, the story changes. When Tesla's inclusion in the S&P 500 was announced in November 2020, the stock jumped more than 13% after hours, pricing the expectation in advance. That much is reflexivity: money that forecast what was coming bought first.

The index-as-device went to work in earnest after that. On 21 December, the day of the actual inclusion, index funds that replicate the index outright had to buy Tesla mechanically. That forced buying ran to an estimated $80–100 billion. Tesla's market capitalization as reflected in the index around inclusion was put at somewhere in the $400 billions, so roughly a fifth of it—anywhere from a sixth to a quarter, depending on the estimate—moved for the single reason that it was "going into the index." The thermometer had taken on the thermostat's job as well. We covered this facet elsewhere too (Passive Crossed Half the Market. Who Took Over the Judgment?). How far it pushed the price is disputed. What is not disputed is that money on that scale moved for nothing about the company, but because it had made a list. "Only reflects" turns false the moment something rides on the metric.

Taking the Engine Apart — Four Parts and One Vital Point

The process by which measurement swallows its object is a chain of four interlocking parts.

Part 1, commensuration. The work of converting different qualities into magnitudes on a single scale so that they can be compared. Putting a small liberal-arts college and a vast research university on the same line of a university ranking is one such job. This compression always throws something away. There is always a residue the single number cannot hold.

Part 2, coupling the stakes. The wiring that hangs interests on that number. What gets hung is capital, credit, distribution, attention. This is the step of pinning a bonus to the dashboard.

Part 3, reactivity. An actor who knows something is riding on the number optimizes for the visible proxy rather than for the hidden true target. The true target usually cannot be seen directly, and the proxy is right there in front of you as a number.

Part 4, performing. The object is deformed into the shape of the metric. The metric crosses over from writing the world down to weaving it. And at the moment of remeasurement, the proxy has drifted from the reality it was standing in for. A classroom taught only what the test asks raises its scores without raising learning by as much. That drift is Goodhart's Law.

How the Measurement Engine Turns — Four Parts and a Feedback Loop
How the Measurement Engine Turns — Four Parts and a Feedback LoopHidden target(what cannot be seen directly)Commensuration(many qualities into one number)Coupling of stakes(money, credit, attention ride on the number)Reactivity(optimizing the number instead of the target)Performation(the object reshapes itself to the number)Goodhart collapse(number and target come apart)어긋나도 루프는돈다
A compression of the chain the body already builds — all four parts must be present for the engine to turn; remove one and it drops back to a camera. Arrows assert order and dependence, not magnitude.

"When a measure becomes a target, it ceases to be a good measure." That widely quoted sentence was not written by Goodhart. Someone else discovered the same law separately, in educational testing. Even the law of measurement circulates with its own source misattributed — this piece's theme, one more time.

Goodhart's Law · tracing the source · The oft-quoted "When a measure becomes a target, it ceases to be a good measure" was formulated not by Goodhart himself but by the anthropologist Marilyn Strathern (1997). Goodhart's own 1975 wording was "Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes." The sociologist Donald Campbell (1976) independently discovered the same law in educational testing—the more any social indicator is used in social decisions, the more it comes under corruption pressure and the more it distorts the very process it was meant to monitor.

The case that runs all four parts to the end is the law school ranking. Begun in 1987, it weighted reputation 40%, selectivity 25%, employment 20% and faculty resources 15%, and put schools of quite different character on a single line. That is Part 1.

A school's fortunes came to ride on that line. Part 2. Deans then remodeled their schools to fit the weights. Resources were allocated differently, and jobs were redefined. Part 3.

The drift comes after. One dean said the pressure on test scores might force the school to change its mission of lowering the threshold. The school tilted toward scores, and one part of the "good school" the ranking meant to measure was pared back instead. Part 4. And the ranking made itself true besides: of the applicants admitted to two elite schools at once, 80–90% chose the one that ranked ahead by a statistically meaningless margin.

The Black-Scholes model shows only Parts 1 and 4, pulled out of the chain. An option is the right to buy or sell a given stock at a price fixed in advance. What that right is worth had no settled answer for a long time, and the formula that appeared in 1973 answered with a single number. At first it did not fit market prices well. But once traders began pricing and hedging with the model, actual prices converged toward it. From the mid-1970s until just before 1987 the distance between formula and market price kept narrowing, and one economist went so far as to call option-pricing theory "the most successful theory in all of economics."

What is missing is the middle. Traders did not cheat the model; they used it. Part 3—an actor that reads its own score and moves to raise it—is not here. So Black-Scholes is not a demonstration of the whole engine but a case that isolates one thing: a number can make its object.

For the engine to turn, all four parts must be present, and each part demands one condition. A number you can compress into (commensuration), stakes you can hang on it (coupling), an actor that reads its own score (reactivity), an object that can be deformed (performing). Remove one and the engine drops back into a camera.

But one of the four is open far more often than the rest. The nature of the object usually cannot be changed. Nor can you simply refrain from producing numbers. What you hang on the number, though, is the measurer's to choose. Take the bonus off the dashboard and the markings go back to being markings. That is why Part 2, the coupling of stakes, is this engine's vital point—not because it is the most important part, but because it is the one most often left open. Withholding the score from the object opens the same spot. That method, too, is available only to the measuring side.

So the objection—"isn't this just a bad-metric problem; design it better, or bundle several metrics together"—does not touch the vital point. The drift is not a design flaw in a particular metric but a structural property of every proxy under pressure. Every metric is a stand-in that compresses a rich reality into one number, and commensuration always throws something away. Once something rides on it, optimization flows into the gap between what the metric measures and what the metric leaves out. Bundle several metrics and the score-chasing merely moves to the seams of the bundle.

Metrics that break down slowly do exist. Where the true target and the proxy nearly coincide, where the stakes are small, or where faking is expensive. But slow collapse is not neutrality. And showing this claim wrong is simple: name one social indicator that has been used for a long time with something large riding on it. If it has not drifted from what it meant to measure far enough to make itself useless, the claim falls.

What separates camera from engine is not "neutral tool versus abuse" either. It is whether the object is a strategic actor that reads its own measurement with something at stake. Stars and temperatures do not read their own thermometers, so they sit forever, docile, before the camera. This is exactly why Goodhart's Law and the educational version of it we just saw are laws about "social indicators used in social decision-making" rather than about measurement in general. The moment the thing being measured is a person or an institution, neutrality is not something you can protect by fending off abuse; it is a condition that never held in the first place.

One objection remains, running the other way. "Isn't 'engine' an overstatement that smuggles agency into an inanimate tool? Numbers want nothing. People did the wiring, and the number is only the medium."

That is right. And it is also this piece's conclusion. Numbers want nothing. The side that chose what to hang on what is the side that wants. What the word "engine" points to is not the will of the number but the wiring, and wiring usually runs with the wirer's name erased. What the final question of this piece takes aim at is that erased name.

So far we have looked at the chain itself. Now for the three stages it turns on. Same engine, but a different part looms large on each stage.

The Individual — People Optimizing to a Rule That Isn't Even the Formula

Start with the number in the wallet. The American personal credit score FICO was commercialized in 1989. Its original target is a disposition you cannot observe directly: will this person repay? The score is the proxy for it. Once loan approval and interest rates rode on the number, people began managing the score rather than their ability to repay.

Reactivity twists one turn further here. The FICO score is a composite model that combines repayment history, debt load, length of credit history and more. Yet the rule people hold on to is usually a single one: keep credit-card utilization under 30%.

That rule is not FICO's official standard. The company has said outright that the data does not support the folk belief that a score drops once utilization crosses 30%. Its own recommendation is below 10%, it added, and average utilization in the top score tier runs around 7%. How many people live by the rule, neither the company nor we can count. But the company had to deny it publicly, and that alone shows how far the belief has spread. Optimization happens at a spot insulated twice over—once from the true target by the proxy, and again from the proxy by a mistaken folk belief about it.

Not using a card at all is no answer either. At 0% utilization there is no active repayment history, so the data to judge on disappears. Hence the behavior of charging small amounts deliberately for the sake of the score. Individual reactivity is institutionalized this way, into an industry of its own: credit repair.

China's Sesame Credit stretched the list of what hangs on the number much further. Released in 2015 by an Alibaba affiliate, this score places an individual between 350 and 950 and starts granting privileges at 600. And it ties that score to benefits in ordinary life—loans, of course, but also match visibility on dating apps and deposit waivers when traveling. It is a private program, separate from the Chinese government's social credit system. The wider the stakes hung on a number, the more people reweave their lives into the shape the score favors. The driver watching the dashboard, replicated at the scale of a country.

Individuals change themselves to fit the number. One level up, we reach a place where the object that would have to change has not itself set yet.

The Firm — When the Proxy Is Closer to Construction Than to Discovery

An MIT research team measured how closely the ESG ratings of six major rating agencies agree with one another. What they measured is a correlation coefficient—how far two ratings line firms up in the same order, where 1 is the identical line and 0 is no relation at all. The result: 0.54 on average, ranging from 0.38 to 0.71. On the same firm, the agencies line things up only about halfway alike.

The ESG-divergence study · Florian Berg, Julian Kölbel and Roberto Rigobon of MIT, "Aggregate Confusion: The Divergence of ESG Ratings" (2022). Full citation in Sources below.

There is a control. On the same firm, credit ratings correlate across agencies at around 0.92.

Using that contrast costs something up front. Credit ratings run on a common ladder from AAA to D while ESG agencies each build their own scale, so of course the side sharing a ruler comes out closer. Credit rating agencies are registered and supervised, which pushes their methodologies to resemble one another; ESG has little such pressure yet.

There is a more uncomfortable objection. Issuers designed products to fit the rating models, agencies handed out top grades side by side, and those grades missed together. Agreement among raters is not agreement with the truth.

So what the 0.92 supports is not "credit is true and ESG is false." Only that there is one difference between the two objects. Default is a grader that comes late but comes. Coming late is the cost we just paid. Sustainability has no such grader at all.

The skeptical reader pushes back here too. "Isn't that just measurement error from each agency using a different method?" The original study did decompose the divergence: 56% measurement (the difference from scoring the same item with different data), 38% scope (what to put in and what to leave out) and 6% weighting (what to weight more heavily).

Start with weighting. Weighting is the component that most directly measures normative disagreement, and its share stops at 6%. On "what matters more," the agencies barely diverge at all. That is an unfavorable number for the explanation that sustainability diverges because it is a fight over values.

And yet scope splits at 38%. They largely agree on what matters more, and split on what belongs inside sustainability in the first place. That 38% is not a problem of precision. It is a problem of what goes on the list. There is also the explanation that agencies face different clients: a firm out to measure financial risk and one out to measure social impact cannot have the same list. Even so, the one who chose that list is the agency.

The client explanation says why the lists differ. It does not say why those different lists are sold under the single name "sustainability."

Why the decomposition was possible · The decomposition worked only because the team first mapped the six agencies' items onto one common framework. A common description could be built—and even so, what to put inside it remained the agency's choice.

ESG divergence is not proof that sustainability does not exist in the world. It sits somewhere in the middle, where the true target is partly real and partly constructed at the same time. Only, the constructed share is larger than one would think.

Here comes the objection that contract theory—the field dealing with how the party ordering work sets rewards for the party doing it—already has the answer. "Isn't this a renaming of the multitask principal-agent problem? Pile the reward onto the measurable dimension and effort gets distorted on the dimensions that aren't measured. Contract theory proved that long ago."

A precise point. The law school we saw is exactly the distortion that theory predicts, so it would be an overstatement to say this piece has discovered a new phenomenon. What changes is the point of application. Contract theory starts by taking the principal's objective as given from outside. What we want to look at is what comes before that. Where did the objective come from? As we saw, the hand assembling sustainability is the rating agency's.

So the question left over changes. If part of the objective sets together with the measurement, the question moves from "how far does this metric distort the true target?" to "who chose that objective, and how?" What is new is not the question but the place it attaches to. Not the contract, but the objective the contract aims at. Contract theory asks how far the map distorts the territory; we ask who drew the territory's boundary. Cases where the chain runs all the way through, like the law school, are cases contract theory already knows. Where the objective came from shows itself only where the chain has not finished turning—as with ESG.

Greenwashing happens too, of course: chasing the rating, fitting disclosures to the checkboxes. The object deformed into the shape of the metric. But that is the same part as the score management we saw on the previous stage. What this stage reveals with unusual clarity is the step before it—that the object to be measured has not yet set.

On the corporate stage the object to be measured wobbled. One level up, what wobbles is not the object but the ruler.

Culture — The Side That Chooses the Shape

The third stage is the one where the side holding the ruler swaps out the markings. The place the reader lives inside every day: view counts and recommendations.

In 2012, YouTube moved the core signal of its recommendation and search algorithms from "view count" to "watch time." On the surface it was a technical adjustment, shifting the center of gravity among several signals. But as creators rebuilt their videos to fit the new signal, the standard of "good content" split into a before and an after.

A part with a different grain from the earlier cases comes into view here. Before, the object—the individual, the firm—was deformed to fit the metric. Here the measuring side moves the metric's center of gravity itself, according to the shape it wants.

Whether to watch view count or watch time is not a fact already present in the world. It is a choice the platform made, and that choice governs the shape of what people make and watch. The delegation of recommendation (What I'll Want, and Who I Handed It To) and the outsourcing of the short attention span (Who Chose That Hour Last Night?) that we have already covered are, from this angle, a question of who chooses the shape. Shape is not given but chosen, and the right to choose belongs to the measuring side.

Where the Engine Doesn't Fully Turn

The engine is real, but partial. Nail it down as always turning all the way and it becomes an overstatement.

First, the engine sometimes tears down the truth it stands on. There was a technique in use in the 1980s: portfolio insurance, built to sell automatically as prices fell so that losses were capped. It has been suggested that this technique, leaning on Black-Scholes-family calculations, actually amplified the 1987 crash. MacKenzie called this counterperformativity: using a tool in a way that pushes the object away from the tool's own description. At the same time he nailed down that it "cannot be proved, but neither is there a decisive way of showing it played no part."

After that crash, a volatility skew that had not existed before appeared in the options market and did not go away. It is the phenomenon in which the volatility the market assigns curves according to where the strike—the trading price fixed in advance—sits. Volatility is how much the market expects the value to swing. Black-Scholes predicted that line would be flat, but the curve stayed, and the model's prediction no longer held. The engine that had been dragging the market into its own shape had, by that very operation, torn up its own map.

Second, physical reality resists. However strong the narrative, semiconductor inventory that has already been sold does not come back on a story (the memory-cycle piece). The more the object is matter rather than a person or an institution, and the less the object can read its own measurement, the less the chain closes all the way.

Third, cutting the coupling does not send the engine straight back to being a camera. When a deformation has lasted long enough to harden into a norm or an identity, the shape lingers for a while even after the stakes come off. Much as a classroom grown used to studying for the test keeps the habit once the test is gone. The vital point is a lever, but not a fast-acting cure.

These three places do not refute the engine. They only show what happens when a part goes missing or sets hard. What we stand on is not that the engine always turns all the way, but that neutrality is not the default in front of a number with something riding on it.

What Shape It's Making, and Who Chose It

The moment a number with something riding on it is put in front of an object that can read its own score and react, that number stops reflecting the world and starts making it. Goodhart's Law is not an ailment separate from this engine but its output after a full turn—the collapse phase.

StageWhat is measuredThe proxy numberWhat rides on itWhat this stage shows large
The individualDisposition to repayCredit scoreLoan approval and interest rate · with Sesame Credit, deposit waivers and dating-app visibility tooReactivity becomes an industry of its own
The firmSustainabilityESG ratings (six agencies)Disclosure fitted to the rating, and reputationThe proxy number assembles the object differently for each rater
CultureThe goodness of contentView count → watch timeRecommendation and search exposure, that is, distributionThe measuring side swaps the number out
Table sources · Individual = FICO's own materials and the Sesame Credit launch announcement · Firm = the MIT team's ESG-divergence study · Culture = YouTube's official blog. Full citations in Sources below.

On none of the three stages is the hand on the vital point held by the side being measured. Individuals only live to the score; what the score gets hung on is settled by the side that makes and uses it. Firms hold about half of it—they disclose to the rating and negotiate as they go, but what goes on the list is the agency's call. Creators hold nothing at all. The platform chooses what to count in the first place.

This collapse, the engine's output after a full turn, cannot be fixed with a better metric. The drift is not a bug in a bad metric but a property that comes out of every proxy under pressure. What can be acted on is not the metric's sophistication but the vital point. Cut the stakes from the number, or hide the measurement from the object.

The test takes two questions. Is there a real event that will grade this number, and how late does that grader arrive? With no grader, the number leans toward construction rather than discovery. With one that arrives long afterward, as default does, the engine wins that much free time. And even where there is a grader, this test does not answer who wrote the grading sheet.

Once you know the wiring, diagnosing life-by-the-score as personal cunning or vanity misses the mark—someone read the wiring and moved the way the wiring says.

So the question to ask when facing a metric changes. The first question is not "Is this number accurate?" Push the accuracy as far as you like and the engine keeps turning. The real question is this: "What shape is this number giving the world, and who chose that shape?" YouTube's decision to take watch time over view count; the standard of sustainability that each ESG agency assembles differently; the rule by which an index committee decides what goes into the market and what stays out. The claim to neutrality—that the number decided objectively—erases from view the very side that chose the shape.

Measurement does not photograph the world. It pushes the world into a different shape. So what we must ask is not the accuracy of the dashboard's markings but who designed the engine to turn the way it does.

Sources
  1. Donald MacKenzie, An Engine, Not a Camera: How Financial Models Shape Markets (MIT Press, 2006) · "Is Economics Performative? Option Theory and the Construction of Derivatives Markets," Journal of the History of Economic Thought 28(1) — https://gwern.net/doc/economics/2006-mackenzie.pdf
  2. Milton Friedman, Essays in Positive Economics (University of Chicago Press, 1953)
  3. Michel Callon (ed.), The Laws of the Markets (Blackwell, 1998)
  4. J. L. Austin, How to Do Things with Words (Clarendon Press, 1962; lectures 1955)
  5. Wendy Espeland & Michael Sauder, "Rankings and Reactivity: How Public Measures Recreate Social Worlds," American Journal of Sociology 113(1), 2007
  6. Wendy Espeland & Mitchell Stevens, "Commensuration as a Social Process," Annual Review of Sociology 24, 1998
  7. Charles A. E. Goodhart, 1975 conference paper for the RBA (Reserve Bank of Australia), Sydney — Papers in Monetary Economics, Vol. I (Reserve Bank of Australia) · Marilyn Strathern, "'Improving ratings': audit in the British University system," European Review 5(3), 1997, p.308
  8. Donald T. Campbell, "Assessing the Impact of Planned Social Change," Occasional Paper #8 (Western Michigan University, 1976)
  9. Florian Berg, Julian Kölbel & Roberto Rigobon, "Aggregate Confusion: The Divergence of ESG Ratings," Review of Finance 26(6), 2022 — https://q-group.org/resources/Documents/Rigobon_Aggregate%20Confusion%20Paper.pdf
  10. S&P Dow Jones Indices, "Tesla Set to Join the S&P 500" (2020-11-16) · inclusion effective 2020-12-21 — https://press.spglobal.com/2020-11-16-Tesla-Set-to-Join-S-P-500
  11. Morningstar, US Fund Flows — passive assets overtaking active within US equity funds (as of 2019-08)
  12. myFICO (Fair Isaac), "History of the FICO Score" · "What Should My Credit Utilization Ratio Be?" — https://www.myfico.com/credit-education/blog/credit-utilization-be
  13. Ant Financial, Sesame Credit (Zhima Credit) launch announcement (2015-01-28)
  14. YouTube Official Blog, "YouTube analytics now includes time watched" (2012) — https://blog.youtube/news-and-events/youtube-analytics-now-includes-time_11
Analyzed and verified multi-dimensionally with AI; reviewed by the author.