Sarthak Garg

Time-boxing and kill criteria

Kill criteria written in advance are worth more than any retro.

·8 min read·

The hard part of stopping a project is not the decision but the timing. By the time you are ready to admit a bet is not working, the team has already spent three months on it, the demo is half-built, and pulling the plug now means standing in front of everyone and calling that time wasted. So instead of pulling it you ask for two more weeks, and then for two more.

A retro cannot fix this, because a retro happens after, when the spend is already sunk and the only honest verdict is the one that costs you the most to say out loud. The version of you sitting in that room is the worst possible judge of whether to continue, because that person is defending a choice they already made.

So move the decision earlier, to the one moment you are not yet invested: before the bet starts. A kill criterion is a line you write down in advance that tells a later, more compromised version of you when to stop, and it beats any retro for the unglamorous reason that it was written by someone with nothing yet to defend.

Write the line before the bet starts

Once you have publicly committed to a project and poured effort into it, every additional week makes stopping harder, because admitting the bet failed now means admitting the last month was also a mistake, and the cheaper move, emotionally, is always to keep going and hope. People call this escalation of commitment, and it is not a character flaw so much as the default behaviour of anyone with skin in a decision.

A kill criterion is a pre-commitment aimed straight at that future self. You write it while you still have no sunk cost and no team's morale riding on the answer, so that when the pressure arrives you are no longer asking whether you want to stop, which has no clean answer, but whether the thing you already agreed would mean stop has happened, which does.

Separate the clock from the verdict

Most people collapse two different things into the phrase "time-box," and that is exactly where it goes wrong.

A time-box is a date, and its job is to force you to look: when it expires, the team has to stop, look at what it has, and decide. The box itself never fires the kill. All it does is open the conversation.

A kill criterion is the verdict you reach in that conversation. The clock exists to make you look, and the criterion tells you what you are looking for when you do.

Keep the two separate and a missed box becomes a real checkpoint. Blur them and you get the failure I see most: the time-box that rolls forward.

Say a team caps a research spike at five days. Day five arrives, the work is not done, and someone says we are close, give it two more, which sounds reasonable and on day seven sounds reasonable again. Nobody set a bar for what an extension has to prove, and nobody in the room pays anything today for granting one, so the box never bites. It has quietly turned into an open-ended commitment with a calendar invite attached. An extendable box is fine, but only if each extension has to clear a fresh bar: what specifically will be true next time that is not true now.

Commit the box and the scope, not the number

This is where I part ways with the standard advice, which says to name a specific number or accept that your criterion is useless: hit forty percent activation or kill it, reach a thousand weekly users or kill it.

I think the number is the least important part. A hard threshold on a metric you have never measured before is a guess dressed up as rigour, and the two things worth pre-committing are the time-box and the scope.

We built an AI reconciliation engine last year. The plan was deliberately bounded: a v1, then one iteration to improve it, rolled out to a small cohort, with the adoption funnel monitored. We never set a specific adoption number. The rule was simpler: two iterations, this cohort, and if adoption does not meet expectations by the end, we pause.

That criterion was obeyable without a number because everything around it was hard: the iterations were capped, the cohort was small, and the date was real. A monitored funnel plus honest judgment is a perfectly good stopping rule, as long as it sits inside a box and a scope that do not move.

Compare it to the criterion that does nothing: "we'll stop if it's not working." There is no date that forces the look, no cap on what gets spent first, and "working" gets judged in the moment by the people most invested in the answer. It is not a stopping rule. It is a feeling, written down.

Cap the spend before the signal

The cohort cap matters more than the threshold, and the reconciliation engine is also where I learned the limit of my own discipline.

The criterion did its job, but it could not fix the fact that v1 itself was big. It took a long time and a lot of effort to build before any adoption signal came back, and when the signal came, the accuracy was average and the adoption was poor. The time-box bounded the iterations, but it could not retroactively shrink the v1 we had already committed to.

The cap you most need is on the spend before any signal comes back. The small cohort saved us, because it kept the eventual climb-down cheap, but if the first version is too large to reach a signal cheaply, the kill always arrives late and expensive, no matter how clean the criterion reads on paper. Scope the first look so the bet can fail before it gets big.

Make it a contract everyone signed

A kill criterion that exists only in the manager's head is a private opinion waiting to be sprung on people later.

The engineers building the reconciliation engine knew from the start that it was a cohort experiment that could be paused, and that mattered more than I expected. When the pause came, the team had already agreed to the shape of that decision months earlier, so it did not land as a verdict handed down on their work.

The version that breaks people runs the other way: the team pours months in believing they are building the real thing, and learns at the funeral that it was "just an experiment all along." That sentence, said for the first time on the day you kill the work, is where the demoralization comes from, and it is why the criterion is as much an emotional contract as a decision rule. Everyone executing has to know upfront that the work is killable and on what line, so that when the kill comes it is a decision they agreed to months ago instead of something done to them.

Write the way back

The fear that keeps good criteria from being obeyed is the false negative, the bet you kill that would have been a winner. It is a real fear, and the answer to it is to make the kill reversible rather than to obey the criterion less.

A kill does not have to mean forever. When we paused the reconciliation engine it went to the backlog rather than to the grave, with a note about what would bring it back: other priorities clearing, or a cheaper path to the accuracy it needed. We are planning to pick it up again.

So write two lines: the stop condition, and the condition that would revive it. What would have to become true for this to come back off the backlog? Answer that in advance and the kill stops feeling like a verdict on the idea, because nobody is being told the work was worthless, only that it does not fit now and what would have to change for it to fit.

When not to obey the line

A criterion you obey blindly is its own trap, so be honest about when the line is wrong.

Kill on outcome signal, and never on difficulty signal. A bet that is hard and slow is not the same as a bet that is failing, and plenty of work that turned out to matter was painful in the middle. We paused the reconciliation engine over its poor adoption, and the difficulty of building v1 never entered that argument. Keep going when the outcome feedback is small but consistently improving, because it points the right way, just slowly. Airbnb is the example people reach for here: investors passed on it, the premise of paying to sleep on a stranger's air mattress sounded absurd, and for a long time the numbers were tiny. But the people who tried it came back, and repeat use is outcome feedback, however small the numbers around it. Kill a bet when the outcome bar you pre-agreed to is missed and "it was difficult" is doing all the work in the argument to keep going.

A criterion worth trusting points at the outcome. The moment to override one is when you notice that the thing you have been calling data is your own exhaustion, which is easy to write down here and harder to notice in the middle of a hard quarter.