What to Build When AI Can Build Anything (2027 Guide)
A step-by-step keep/kill method for B2B product teams: build only what named customers validated, ship to them first, and delete what they don't use.
Your team can now ship ten things a week. AI-assisted engineering cleared the capacity bottleneck, and the backlog has become a decision problem. The question is no longer "can we build it?" It is "should this exist, and for whom?" This guide gives product managers, CS leads, and founders a repeatable keep/kill method you can run in a spreadsheet starting today. At BuildBetter, we work with B2B product teams that face this every week. Below you'll find the step-by-step process, a worked example, and the mistakes that most often break it.
What 'deciding what to build' means when AI does the building
Deciding what to build means choosing which requests get built, which customers they ship to first, and which shipped features stay or get deleted. Each of those decisions rests on what named customers actually do, not on what gets proposed in internal meetings.
That definition covers three decisions most teams treat separately: intake, rollout, and deprecation. When engineering was the bottleneck, intake did most of the work, because you could only build a few things and the rest died in the backlog. Now almost everything can get built, so rollout and deprecation carry the weight.
The cost of software has moved. Writing the code is cheap. Living with it is not. Every shipped feature has to be tested, secured, documented, supported, explained during onboarding, and kept compatible with whatever ships next. That cost recurs whether or not anyone uses the feature. An unused feature is no longer a one-time sunk cost. It is a permanent tax on product surface area.
The short answer:
Build only what a named customer asked for and validated. Ship it to those customers first. Keep what they use and delete the rest on a fixed schedule.
One rule sits underneath the whole method: customer truth is the tiebreaker. When sales, product, and engineering disagree about what to build or keep, the evidence from named customers decides. It doesn't matter who argues loudest, who has the most seniority, or which department owns the quarter's target.
Why this is harder now: what breaks when shipping is cheap
Cheap shipping removes the natural filter that used to keep weak ideas out of the product. When every feature took a sprint, scarcity forced prioritization. Without that scarcity, several failure modes speed up at once.
- Velocity stops meaning anything. Story points and shipped counts measure motion, not progress. In a 2023 controlled study, developers using GitHub Copilot finished a defined coding task 55.8% faster. A 2025 METR study found experienced open-source developers using AI tools took 19% longer, even though they believed afterward they had been about 20% faster. Perceived speed and real progress diverge. Shipping faster in the wrong direction gets you lost sooner.
- Nothing gets deleted. Teams add features at AI speed and remove them at committee speed. Feature bloat compounds, and every new feature competes for attention with forty ignored ones. A widely cited 2019 feature adoption report estimated that 80% of features in the average software product are rarely or never used. The figure is directional, but it matches what most PMs see in their own analytics.
- Requests lose their owners. A ticket that says "customers want X" with no names attached can't be validated, followed up on, or closed.
- Anonymous vote counts mislead. A hundred anonymous upvotes tell you less than one named customer explaining the problem they hit last Tuesday.
- Stated preference drowns out revealed preference. Customers say they want things they won't use. Exit surveys capture polite reasons such as "price" or "missing feature" instead of the behavior that led to churn.
- Roadmap theater gets worse. Teams commit months of features to customers who never asked, because building them feels nearly free.
- Sales promises disappear. Commitments made on calls never reach the product team. The wrong things get built, and the promised things get forgotten.
Experimentation research points the same way. Ron Kohavi's work at Microsoft found that only about one-third of tested ideas improved the metric they were designed to improve, and in heavily optimized products like Bing the rate fell to roughly 10–20%. If most ideas from experienced teams don't produce value, deletion should be the default.
The keep/kill method: a step-by-step process you can run in a spreadsheet today
The keep/kill method is a nine-step loop that turns customer requests into shipped, measured, and either kept or deleted features. It comes from Customer-Led Development: How to Build What Your Customers Want When AI Can Build Anything by Spencer Shulem. It needs no budget and no tool, only a shared spreadsheet and a fixed calendar slot.
Step 0: Set the tiebreaker
Agree in writing that customer truth resolves conflicts. A team can hold only one decision philosophy. "Sometimes product-led, sometimes customer-led" means the loudest department wins. Write the rule down, get leadership to sign off, and point to it when disagreements come up.
Step 1: Apply the name gate
Every row in the spreadsheet needs at least one named customer, meaning an account and a person, plus the problem in that person's own words. No name, no build. This one rule filters out most internal pet projects and anonymous noise.
Step 2: Validate before building
Spend a few days confirming the problem with the named customers through a 15-minute call or a written question. Customers usually describe solutions ("add a CSV export") rather than problems ("I rebuild this report by hand every Monday for my CFO"). Validate the problem, and let the PM's judgment shape the solution. Customer-led is not customer-dictated. Following Rob Fitzpatrick's The Mom Test, ask about past behavior ("When did this last happen? What did you do?") rather than hypothetical intent ("Would you use this?").
Step 3: Log promises
If sales or CS committed to something, record who promised it, to whom, and by when, in a column the product team can see. Promises hidden in call notes are the most expensive kind of request, because they carry relationship risk and nobody tracks them.
Step 4: Ship to the customers who asked, first
Release to the named requesters before general availability through a feature flag, a beta toggle, or a direct invite. This isolates their usage so you can measure it on its own instead of averaging it into the whole customer base.
Step 5: Close the loop
Tell each named customer the feature shipped and how to use it. Closing the feedback loop is part of the measurement. A feature the requester never heard about can't be judged fairly on usage.
Step 6: Watch behavior, not opinions
Before shipping, choose a review window, for example two to four weeks. When it ends, check what the named customers actually did. Use your existing outcome metrics as lenses: retention, engagement, conversion, and expansion. This is revealed preference, an economics idea Paul Samuelson introduced in 1938: preferences show up in choices, not in stated opinions.
Step 7: Make the keep/kill call on a fixed cadence
In a recurring review, keep what the named customers use, delete what they don't, and record the reason in one line. Deletion is the default. Keeping has to be earned by evidence.
Step 8: Measure the system correctly
Track validated outcomes (features kept because requesters used them) and closed loops per cycle. Stop reporting shipped count or velocity as success.
Spreadsheet template
| Column | What goes in it |
|---|---|
| Request | Short label for the request |
| Named customer(s) | Account + person for each requester |
| Problem in their words | Direct quote, not a paraphrase |
| Source | Call, ticket, Slack thread, email |
| Promise? | Who promised, to whom, due date |
| Validated? | Y/N + date |
| Shipped to requesters | Date |
| Loop closed | Date each requester was told |
| Usage at review | What requesters did during the window |
| Metric moved | Retention, engagement, conversion, or expansion |
| Decision | Keep / Kill / Extend once |
| Reason | One line |
Decision rules
- Keep if the named requesters use it repeatedly, or if a chosen outcome metric moves for their accounts.
- Kill if the requesters don't use it within the window.
- Extend once, and only once, if the loop was never closed.
- Kill if no named customer can be found at review time.
Worked example: 50 requests shipped, 10 kept, 40 deleted
This example shows how one quarter of keep/kill decisions plays out in practice. It is illustrative, not a customer case study, but the proportions resemble what teams report when they first run the method.
Setup: a B2B team with AI-assisted engineering ships 50 small requests over one quarter. Each one goes only to the customers who asked, and each has a three-week review window set before release.
Kept (10)
The named requesters used each of these on a recurring basis. Three examples:
- XLSX export with preserved formulas: requested by four named accounts. All four used it weekly. One expanded seats during the same period. Decision: keep and release generally.
- Bulk-reassign for open tasks: requested by two ops leads. Both used it daily during a team restructure, and support tickets on the topic stopped. Decision: keep.
- Saved filter for overdue renewals: requested by one CS director. She used it every morning and shared it with five teammates. Decision: keep and promote in onboarding.
Deleted, never opened by requesters (18)
The loop was closed and the requesters were told, but no usage followed. This is the gap between stated and revealed preference. Decision: delete, and send each requester a one-line note explaining it's being removed and inviting them to reply if that's a mistake.
Deleted, tried once and abandoned (12)
Each got a first use and no return. The underlying problem may be real while the solution is wrong. The follow-up question for each customer: "You tried it once. What were you hoping it would do that it didn't?" Decision: delete and log the problem for revalidation.
Deleted, no named customer at review (6)
These came from internal ideas or anonymous votes that slipped past the name gate because each "only took an afternoon." Decision: delete and tighten Step 1.
Deleted, promise-driven but unused (4)
Sales had promised these. In one case the PM called the account. The buyer explained that the workflow the feature supported had moved to another team six weeks earlier. A promise the customer no longer needs is information, not failure. Decision: delete, update the promise log, and tell the AE why so the account plan reflects the change.
| Category | Count | Evidence | Decision | Follow-up action |
|---|---|---|---|---|
| Used repeatedly | 10 | Recurring usage by requesters; one seat expansion | Keep | Release generally |
| Never opened | 18 | Loop closed, zero usage | Kill | One-line note to each requester |
| Tried once | 12 | Single use, no return | Kill | Revalidate the problem |
| No named customer | 6 | Internal or anonymous origin | Kill | Tighten the name gate |
| Promise, unused | 4 | Account confirmed need had changed | Kill | Update promise log, brief AE |
Takeaway: measured by velocity, the quarter was 50 features shipped. Measured by validated outcomes, it was 10 features kept, loops closed on all 50, and 40 features removed before they could become permanent support and documentation load. For next quarter, the team made the name gate strict for anything under a day of build time. Those small requests were the ones slipping through.
What to keep, replace, add, and kill in your current process
The keep/kill method sits on top of your existing practices. It does not replace them. You don't have to tear down product-led growth, a sales-led motion, or Agile ceremonies. You change what feeds them and how you judge their output.
| Practice | Verdict | What changes |
|---|---|---|
| Sprints and standups | Keep and upgrade | Add customer names, validation status, and loop closure to every item |
| PLG motion | Keep | Add champion identification so self-serve signups have names |
| Sales team and land-and-expand | Keep | Add promise tracking visible to product |
| Decision philosophy | Replace | Customer truth becomes the tiebreaker |
| Roadmap theater | Kill | Commit only to validated, named requests |
| Anonymous voting board | Kill | Replace with named requests and the problem in the customer's words |
| Exit surveys as churn analysis | Kill | Review the last ninety days of account behavior instead |
| Velocity as a success metric | Kill | Track validated outcomes and closed loops |
The fastest upgrade is the standup. Replace "what did you do yesterday" with three questions:
- Which promises are due this week?
- What is awaiting validation?
- Which loops are still open?
These keep named customers in the room every day without adding a meeting. They also echo Teresa Torres's continuous discovery practice of weekly customer touchpoints, applied to delivery rather than only to discovery.
Common mistakes that break keep/kill decisions
Most keep/kill programs fail because the team skips a step, not because the method is wrong. These are the failure modes we see most often.
- Choosing the success metric after the data comes in. If you pick the metric once you see the numbers, every feature looks like a keeper. Write the metric and window in the spreadsheet before shipping.
- Shipping to everyone at once. The requesters' usage gets lost in general noise, and you can't tell whether the people who asked actually adopted it.
- Skipping loop closure, then deleting features the requester never knew existed. This produces false kills and angry customers.
- Having no fixed cadence. This is the most common failure. The review slips one week, then three, and nothing gets deleted. Put a recurring review on the calendar and treat it like a board meeting.
- Treating customer requests as specs. The PM still owns the solution. Validate the problem, then design.
- Letting anonymous votes or internal enthusiasm bypass the name gate because "it only took an afternoon to build." Build time is no longer the cost. Maintenance is.
- Deleting silently. Removing a feature without telling the customers who asked erodes the trust that makes them willing to give feedback next time.
Some failures no tool fixes: leadership that won't allow anything to be deleted, PMs who won't call customers to validate, and sales leaders who refuse to log promises. Software, including BuildBetter, can surface the evidence. It cannot make the team act on it. If the tiebreaker from Step 0 isn't real, the rest of the method won't hold.
When a spreadsheet stops working, and which tools help
The keep/kill method needs no tool. Tooling becomes useful once the evidence for each decision lives in conversations rather than in the spreadsheet. Rough signals that you've crossed that line:
- Tracing a row back to who asked and why takes more than a few minutes.
- Several people are logging requests from calls, each in their own format.
- Promises get made in meetings nobody transcribes.
- Feedback volume passes roughly a couple hundred pieces a month. This is a rule of thumb, not a statistic.
1. BuildBetter
BuildBetter captures calls through no-bot local recording, a bot recorder, or mobile, and combines them with Slack threads, support tickets, and surveys through 100+ integrations. It ties named customers and their exact words to each request through Signals, tracks commitments with Tracked Objects, creates Linear and Jira issues with full context through Tickets, and drafts loop-closure emails when something ships. Because it combines external customer feedback with internal team activity, the promise made on a sales call and the ticket in your backlog end up linked. It works best when your evidence lives in conversations.
2. A shared spreadsheet
Free, and enough for a small team with a single intake owner. It breaks when sources multiply and nobody can keep the name and quote columns current.
3. Airtable
Relational records link customers, requests, and decisions, and you can build review views for the keep/kill meeting. Evidence capture is still manual.
4. Coda
A doc-and-table hybrid that suits teams who want the review ritual and the decision log in one place. Evidence capture is still manual.
A caveat: BuildBetter is not built for enterprise-scale survey distribution or for use as a dedicated research repository. If your evidence comes mainly from large survey programs, you may need a survey-specific tool alongside it.
FAQ
How do you decide what to build when AI can build anything?
Build only what named customers asked for and validated. Ship it to those customers first, tell them it exists, watch whether they use it during a window you set in advance, and keep or delete it at a fixed review.
When should you kill a feature?
Kill it when the customers who asked for it don't use it within a review window you set before shipping, after you have confirmed they were told it exists. Also kill it if no named customer can be found for it at review time.
Are feature voting boards a good way to prioritize?
No. Votes without names can't be validated, followed up on, or closed. One named customer explaining their problem is worth more than a hundred anonymous upvotes.
Is velocity still a useful metric with AI-assisted development?
Not as a success metric. Velocity measures motion, not progress, and research such as METR's 2025 study shows that perceived speed and real productivity can diverge. Track validated outcomes and closed loops per cycle instead.
Does customer-led mean building whatever customers ask for?
No. Customer truth informs the decision, and the PM's judgment shapes the solution. Customers usually describe solutions, so validate the underlying problem before deciding what to build.
Do you need software to run keep/kill decisions?
No. A shared spreadsheet and a fixed review slot are enough to start. Tools help once the evidence is spread across calls, Slack, and tickets.
Sources and further reading
- Spencer Shulem, Customer-Led Development: How to Build What Your Customers Want When AI Can Build Anything, Chapter 19, "What to keep, what to kill" (book).
- Ron Kohavi, Diane Tang, Ya Xu, Trustworthy Online Controlled Experiments (Cambridge University Press, 2020).
- Peng, Kalliamvakou, Cihon, Demirer, "The Impact of AI on Developer Productivity: Evidence from GitHub Copilot" (2023).
- METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (July 2025).
- Rob Fitzpatrick, The Mom Test.
- Teresa Torres, Continuous Discovery Habits.
- Revealed preference (Wikipedia): https://en.wikipedia.org/wiki/Revealed_preference
- Agile software development (Wikipedia): https://en.wikipedia.org/wiki/Agile_software_development
- BuildBetter: https://buildbetter.ai
Make churn optional.
Your team can build anything now. The hard part is knowing which requests have a named customer behind them, which promises are due, and which shipped features nobody touched. BuildBetter connects every call, ticket, and Slack thread to the customers and commitments behind them, so your keep/kill review runs on evidence instead of memory.
Make churn optional. Book a demo →
Further reading: Customer-Led Development
The method on this page comes from Customer-Led Development: How to Build What Your Customers Want When AI Can Build Anything by Spencer Shulem, BuildBetter's founder, drawn from more than 2,000 conversations with product leaders. It lays out the full system: how customer truth flows through a company, how every feature traces to a named customer, and how loops get closed.
Start with the overview: What is customer-led development?