Sarthak Garg

Framing bets so they can fail cheaply

Design bets where failure is legible, survivable, and a source of learning.

·11 min read·

Picture the typical EM's quarter. It opens with ten bets in flight; eight get built, two slip, but none of the eight produces a clear answer. The PM is already lining up ten more ideas for next quarter, also small and unconnected, and the retro is adding a new check because last sprint's spec arrived late. The team is moving, busy, exhausted, and going nowhere in particular.

The lean startup method was supposed to fix this: build, measure, learn, with cheap experiments that kill what does not work. The framing has been around for over a decade, and most teams quote it back fluently. Somewhere on the way down from the slide deck to the sprint, though, the essence got sanded off. What survived was "be fast," and what got lost was the part about pointing the bets at a goal and treating failure as information.

I think of each bet as an arrow with a direction and a size. One enormous arrow means the whole quarter rides on a single shot, and a hundred small arrows pointing every which way means motion without progress. What compounds is a small number of small arrows all pointed at the same goal, with the direction nudged based on what each one returned.

You can only nudge if failure is cheap, because a bet that wrecks the quarter when it misses gets survived rather than adjusted. So cheap failure is design work, done before the bet starts, and the slogan never specifies what that work is. You design the bet so that a no answer is survivable and clear, and so that the team walks out with sharper direction even if the bet itself did not work.

What follows is the playbook I run for that: seven steps for designing and running a bet, then a closing section on how the playbook gets misused.

Set the goal before you size bets

Pick the goal first, decide how many bets the quarter can carry, then write them out.

Most teams do this in reverse. A PM walks in with ten product ideas they have conviction in, the team gets excited and builds them, and none of the ten are aimed at a stated goal, so nothing learned in one bet feeds the next. The quarter ends with a lot shipped and the business number still flat, the team is told to be more product-led, and the cycle repeats. It repeats because nobody inside it did anything that looks like a mistake: the PM brought ideas, the engineers shipped them, and the quarter reads as a productive one right up until someone asks what any of it changed.

The version I run is unglamorous: name the goal in one sentence, whether that is average revenue per user, activation in the first week, or deflecting a specific kind of support ticket. Then cut the bet list to five things that could plausibly change that number, ranked by which has the strongest hypothesis rather than by who is most excited. The capacity you did not commit is not going spare, because the doubling-down step further down runs on it.

This step comes before portfolio thinking, which is about mixing safer work with riskier bets across the quarter. This one points each uncertain bet at the same goal, so the learning from one feeds the design of the next.

Write the hypothesis, including how it dies

Most specs read like contracts: we are building X to achieve Y, ship date Z. Success is implicit, the team works as if the feature is sure to work, and when it does not, the failure feels like betrayal instead of information.

A bet meant to fail cheaply is written differently. Before the team starts, four things go on paper:

  • The change being made.
  • The specific ways it can fail.
  • The reason it is worth trying anyway.
  • The thing the team hopes to learn whether it works or not.

The fourth line does the real testing. If you cannot name what a no answer would teach you, the bet is only designed to succeed, and a no answer will arrive as an embarrassment rather than a result.

Individual feature bets in startups succeed around one time in four, and most specs are written as if the rate were four in four. That single mismatch is behind a great deal of the demoralisation and process-bloat that follows a failed bet: the team is shocked by the no answer, and the answer is shocking only because the spec lied about the odds.

The named ways a bet can fail come out of failure rehearsals. Run one before you write the hypothesis, once the bet is big enough to be worth the time.

Cap the count and double down on what works

The first instinct, given a long list of plausible bets, is to run them all in parallel. Resist it, and halve the count you first wanted.

With fewer bets running, you can see what each one is doing. Ten bets in flight make every weak result look like every other weak result, while five leave enough attention per bet to spot what is moving and act on it early. Kill the duds inside the first two weeks, then take the freed capacity and double down on the one or two that worked: tighten the scope, expand the rollout, or start a second bet that builds on what the first one taught you. The compounding happens in that doubling-down.

Most teams treat this part as an accident. They start the quarter committed to ten bets and finish it committed to the same ten, and the only variation is which ones got the most love, which is what killing a bet looks like when nobody wants to say the word out loud. Plan for the move from week one, so that shifting capacity into the winners is the default rather than the exception. Per-bet stopping rules belong in time-boxing and kill criteria; the discipline here is team-wide: cap the count, and move capacity to the winners.

Strip the MVP to must-haves

The MVP, or minimum viable product, exists to test the hypothesis and nothing else.

A team that builds a "minimum viable" version that is somehow polished and true to the designer's full vision has really built the second version of a product it already decided was going to win. Good-to-haves in the first cut are a tell: the builder is trying to validate their bias rather than the hypothesis.

Strip the bet to the smallest version that could plausibly test the hypothesis you wrote earlier. If the hypothesis is that users will pay for X if you offer it, the first cut is a paywall and a charge, and the beautifully designed checkout comes later, assuming there is a later. If the answer comes back yes, you build the rest; if it comes back no, you have not sunk three weeks of design on something you are about to throw away.

Set the rhythm: checkpoints in, mid-sprint changes out

Once the bets are running, two rules keep the work on track without freezing it.

The first is a scheduled checkpoint: the team gathers on a fixed rhythm, looks at what each bet has produced so far, and makes small corrections. The corrections stay small, because the checkpoint exists to catch drift early rather than to redirect the bet.

The second is a freeze on changes to the plan mid-sprint, unless the issue is a true emergency. The rule looks restrictive until you watch a team run without it: every weak early result becomes a debate, every debate becomes a scope change, and no bet survives long enough to produce a clean answer. With the freeze in place, the team can trust that the plan holds until the next checkpoint, which is when the bet gets evaluated.

I ran the pair on a team I moved off the standard engineering-and-product split into a builder-mode shape, where engineers carried more end-to-end ownership. Part of the reason was the AI era itself: models now fill more of the context-gathering work the spec used to do. Before we started, I told the team explicitly that the first few sprints would surface issues, which framed the risk as intentional before anyone had hit an issue. Inside the sprints I held the no-changes-mid-sprint line and did not adjust the shape on every wobble, and at the checkpoints we looked at the inputs together, picked the small adjustments worth making, and left the early results alone in between. That rhythm is the only reason the bet could be evaluated at all.

After failure, classify before correcting

The most expensive thing an EM can do after a failed bet is to add a new process check.

The end-of-first-sprint retro for that builder-mode transition produced a stack of inputs, one of them sharp. Someone said specs arrived late and were half-baked, and proposed adding a check to make sure they were early and complete from then on. The team's instinct was to add the check to the working agreement and move on.

I held off. The first sprint of a new working shape is exactly when first-attempt failures are expected, so the risk was intentional and the failure was anticipated: the input was a known cost of running the bet rather than a flaw the bet had revealed. The right response was to log the learning and revisit only if the pattern kept showing up after the team should have settled in. I took the same position on several similar inputs from that retro.

The classification comes down to two questions, asked before any process change. Did this fail because the risk was intentional and the answer was no? Log the learning and resist changing anything. Did this fail because the design was bad, with scope too large, no ways to fail named, or a vision built instead of an MVP? Fix the design rather than the process.

Skip the classification and every failure generates a new check. Each check was a rational answer to one failure, and no retro ever convenes to remove one, so the pile only grows: six months in, no one can ship without three approvals, and no one remembers which approval was supposed to prevent which failure. The team pays in risk aversion, accumulated one well-meaning retro at a time. The retro itself has its own craft, covered in postmortems that change behaviour; the classification is the filter you run before it, deciding whether the retro should produce learning or a process change.

Back conviction when the room earns it

The playbook so far has been about breaking big bets into small reversible shots, and the exception matters.

Sometimes one person on the team has very high conviction in a single big-arrow bet, they have earned that conviction, and the room is smart enough to back them. Slicing that bet into many small reversible shots is the wrong call: it dilutes their shot and treats the playbook as dogma rather than design.

The case I learnt this on was a new product the team created. One builder had very high conviction in the direction, the room was smart, the builder had a track record, and so we backed the big-arrow bet whole. The reverse call, mechanically slicing it down because the playbook said so, would have killed the bet on rhythm alone.

This is the situation disagreeing and committing exists for. If you are in the room and you disagree, you record the disagreement, you commit to the call, and you give the builder real cover to run the shot. The room and the person carrying the bet are doing the work here, and the framework substitutes for neither: in a less smart room, with less earned conviction, the same call fails, and the playbook does not protect you from that.

When this gets it wrong

The playbook itself goes wrong in two ways.

The first is applying small-bets discipline mechanically to a conviction shot: the bet gets sliced into reversible cuts because that is how we do things here, and the person carrying it never gets to run the real bet. The exception step above is the corrective, and it holds only if you remember that the playbook is design work rather than dogma.

The second is tightening the rhythm rules until they suffocate the work: daily course corrections, a check added for every retro input. The rules were supposed to protect the team from thrashing on early results, and pushed too hard they produce the opposite failure, where the team thrashes on the rules instead of on the work. If a rule starts protecting the process rather than the bet, drop it.

One last thing on the culture argument. The fail-fast slogan attracts two standard critiques, and both are fair: people learn less from failure than they think, and the office ritual of celebrating failure is mostly theatre. Both are aimed at the slogan, though, and the practice underneath it does not need the slogan defended. Do the design work the slogan never specified, and the culture argument goes quiet on its own: people discuss failure comfortably when it genuinely carries information, and defensively when the design forced it to be embarrassing.