What not to measure: vanity and surveillance
LOC, PR count, keystroke telemetry. Refusing to measure is a leadership act.
Engineering organizations accumulate numbers that count activity rather than results: lines of code written, commits pushed, pull requests merged, story points burned down. The standard advice about these numbers is settled and correct, and I follow it. Activity counts measure volume while the business cares about value, and once one of them touches a performance conversation, people optimize the number instead of the work. Goodhart's law gives the mechanism its name: attach stakes to a proxy and the proxy comes loose from the thing it was standing in for. Reward line counts and you get padded diffs and quietly skipped refactors; reward pull request counts and you get work sliced thin to inflate the total. Designing metrics that resist that gaming is its own discipline, and the advice concludes, correctly, that you should measure outcomes instead.
The standard advice does not cover what happens next: the forbidden numbers come back. The quarterly review template arrives with a per-engineer column already in it, and the project tracker renders a per-person velocity chart by default. The director who agreed with the Goodhart argument in the spring asks for commit counts in the autumn, because budget season started and every other function brought per-head numbers. Not measuring, it turns out, is not a conclusion you reach once but a set of refusals you make repeatedly, each with its own trigger and its own way of failing, and holding them takes as much design effort as the measuring itself.
Keep activity counts off individual names
A team's cycle time trend, how long work takes to get from started to shipped, or its weekly throughput tells you something about the system the team works inside, and the flow metrics worth tracking cover that list. The per-person slice of the same data tells you almost nothing about any individual, because output volume depends on what each engineer was assigned, how tangled the code they touched was, and how much of their week went to reviewing and unblocking other people. So I hold a simple rule: activity metrics live at team level or they die, and individual engineers get evaluated on outcomes and judgment and never on volume.
A leader prepping for a quarterly business review opens the shared template and finds a column headed pull requests per engineer. Filling it in takes two minutes, and deleting it takes a small act of will, because the column implies that someone upstairs expects it. The leader deletes it anyway and replaces it with three bullets on what the team shipped and what changed for customers.
Delete the column and offer nothing in its place, though, and the refusal reads upward as evasion and invites your leadership to collect the number without you. The substitute matters as much as the refusal, and a later move, bringing a better number upward, is about building it.
Say no to surveillance as policy
The second refusal is about a different kind of number altogether. Keystroke loggers, screen recorders, presence trackers, and activity scores are not badly designed metrics that a better dashboard could fix. They measure the person rather than the work, and they change the employment relationship the day they are switched on. An engineer who knows their keystrokes are being counted starts performing activity, spacing commits through the day and keeping an editor window warm, and none of that is engineering. Meanwhile the real work of the job, which on a hard day looks like an hour of reading and twenty minutes of typing, becomes something to feel anxious about. Teams that make this trade tend to get slower, and they pay in trust.
This plays out publicly on a regular cycle. A large company rolls out workforce telemetry, employees discover the dashboard, the internal forums revolt, senior engineers start interviewing elsewhere, and the program is quietly scaled back with a note about good intentions. The people who leave first are the ones with the most options.
I hold this refusal as policy rather than preference, stated to the team: we do not monitor individual activity telemetry. Framed as my personal discomfort, it would last until the next reorg. Stated as policy, it cannot truly bind whoever sits in the chair after me, but it forces them to overturn a published commitment out loud instead of letting a preference quietly lapse.
Publish what you do watch
That policy has a boundary problem, because I do watch people. I watch review latency creep upward and calendars fill with meetings, and I notice when a normally vocal engineer goes quiet in design discussions; watching for leading indicators of burnout demands exactly that kind of attention. If keystroke telemetry is banned, something has to explain why this is allowed. I use three tests to stay on the right side of the line:
- Work or person. Does the signal describe the work or the person? A pull request that sat unreviewed for four days describes the work; an engineer's active hours describe the person.
- Aggregate or individual. Is it read in aggregate first, with individual attention arriving only when someone needs help rather than evaluation?
- Out loud. Would I be comfortable telling the team what I watch and why?
Anything that fails a test stays out, and everything that passes gets said out loud. The second test is the one I trust least, because aggregate-first is easy to claim and hard to audit from inside your own head.
The watching is legitimate because the team knows about it. A private spreadsheet of review latencies is surveillance even when the intent behind it is kind, because the people being watched have no idea and no say.
In a team meeting, the leader lists the signals they watch and explains that the purpose is spotting who needs help before they have to ask. Someone asks whether any of it feeds performance ratings, and the leader says no and names what does. The list becomes part of the team's working agreement.
Publishing a list and then extending it without saying so converts observation back into surveillance on the day anyone discovers the difference.
Bring a better number upward
Sooner or later the ask comes from above: rank the engineers, or tell us who the bottom ten percent are. In 2023 McKinsey published a framework claiming that individual developer productivity can be measured, and Kent Beck and Gergely Orosz wrote the rebuttal that became the reference point for the whole debate: the framework counts activity, activity is not productivity, and measuring a surgeon by time spent holding the scalpel tells you nothing about whether the patient recovered. Practitioners largely landed with Beck and Orosz, yet the ask keeps resurfacing in executive rooms at budget time, because it is rarely about the metric itself.
Underneath the ask sits a real question, usually some version of: are we getting value for this headcount, and how would we know? You owe that question an answer, and a lecture on Goodhart's law is not one; it does not survive contact with a budget meeting. Bring a substitute that answers the question better than the forbidden number would have: the outcomes the team shipped and what they changed for the business, plus the team's cycle time trend and its quality record in incidents and defects that reached customers.
The measurement frameworks get the same handling. SPACE, a framework for developer productivity, and DORA, the standard set of delivery metrics, both warn in their own papers against single metrics and against grading individuals, and both are routinely bought as dashboards and deployed as the individual scorecards their authors warned about. When a framework arrives in your inbox as a mandate, the refusal and the substitute work the same way as they do for the raw numbers. When the ask is a ranking of teams rather than people, comparing teams fairly takes over.
Delete the metrics nobody acts on
The moves so far answer other people's asks; this one is hygiene on my own dashboards, where the enemy is convenience rather than pressure. Metrics accumulate because tools collect them for free, and each one costs a little attention. Twice a year, on a calendar reminder because it does not happen otherwise, I run the same test over every chart, the test dashboards that get read is built on: who reads this, and what decision changes when it moves? A metric with no reader and no decision is decoration, and decoration that ranks people is worse than decoration. Everything that fails the test gets deleted, and each deletion gets an owner, someone whose job includes noticing when the chart creeps back.
The chart does creep back, because the tool never stops offering it. A per-person velocity chart gets removed in the spring and returns by the next planning season, because the tracker renders it on every new board and a well-meaning program manager copied last cycle's setup forward.
Delete a number without asking what anxiety it was soothing, and it comes back. If a director kept the chart because it was their only way to feel confident the team was working, the chart was a symptom, and the honest fix is the substitute from the previous move rather than a quieter deletion.
Refuse the AI-era reruns by name
The newest numbers arriving on dashboards are AI adoption rate per engineer, share of code that is AI-generated, assistant hours per day, and leaderboards ranking who accepts the most assistant suggestions. Vendors sell them as productivity benchmarks, and some organizations now tie usage quotas to performance reviews. This is the lines-of-code mistake rerun with fresher numbers, counting activity with an assistant instead of activity in an editor, and it deserves the same refusal, made by name, before the first leaderboard ships.
An organization ties assistant usage rate to reviews, and engineers respond by accepting suggestions they would previously have rejected, then fixing them in a follow-up commit. The acceptance rate climbs and the dashboard reports a triumph, while the review burden grows behind it.
Aggregate adoption data has to survive this refusal, because read at team level it exposes real enablement gaps: a team whose usage is low may have a licensing problem or a codebase the tools handle badly, or it may have nobody with the time to share what works; track that in aggregate and act on it. Ranking individuals on usage compliance crosses the line, and distinguishing human contribution in AI-augmented output takes up the harder question underneath, which is what individual contribution even means once the output is half machine.
What survives the refusals
After all six refusals, a full measurement practice is still standing. Team-level flow and quality metrics survive, and so does published observation. Even the per-person look survives in one narrow form: when I am diagnosing a specific problem with a specific engineer, with their knowledge, for a bounded stretch of time, the same data I refuse to put on a dashboard becomes part of an honest conversation. Each refusal targets a reading of the data more than the data itself, and the readings that go are the ones nobody was told about, pointed at one name at a time.
The refusals are not the passive part of measurement: a team learns what its leader values partly from what gets celebrated and partly from what the leader declines to count and says so. The four classes of signal cover what to measure in place of everything refused here.