Ramp's AI Product Factory: Every Internal Tool Explained (2026)

A tool-by-tool guide to Ramp's AI software factory from Geoff Charles's 2026 talk: Inspect, Glass, Review Buddy, Testo, Gadget, and how to copy each one.

Ramp's AI Product Factory: Every Internal Tool Explained (2026)

In September 2026, Geoff Charles, Chief Product Officer at Ramp, gave a talk at the Lenny and Friends Summit called "The limiting factor — how to design an AI software factory for speed." This guide from the BuildBetter team goes through every internal tool he showed, what each one does, the metric he attached to it, and how you could build the same pattern yourself. His thesis fits in one sentence: AI does not remove the bottleneck, it moves it, and the teams that win are the ones that find the next bottleneck fastest.

That idea echoes Eliyahu Goldratt's Theory of Constraints. A system's throughput is set by its tightest constraint, and when you fix that constraint, a different one takes its place. Charles asked PMs to invest in the factory that builds the product, not only in the product itself. His closing line was blunt: "I want you to copy us." This page is meant to help you do that.

Two disclosures first. Every tool below is internal to Ramp. Ramp sells corporate cards and spend management software, not developer or PM tooling, so none of these tools is for sale. The pattern behind each one is what you can copy. Second, Charles gave his own caveat: "Everything you've seen here is outdated." Read the tools as a snapshot of a system that keeps changing, not as a fixed blueprint.

The structure: five steps (Identify, Define, Build, Coordinate, Improve), seven tools, one summary table, and a "how to copy it" list for each tool.

Every Ramp AI Tool at a Glance (Summary Table)

Ramp's AI software factory has seven tools spread across five steps. Each one exists to relieve the constraint created by speeding up the step before it.

StepToolWhat it doesHeadline numberHow to replicate it
IdentifyCustomer insight agentPulls customer pain from every data source, then clusters and ranks itA 1M-token context window is less than 0.5% of Ramp's Gong transcriptsBuild ETL, vector search and clustering over all feedback sources, or use an off-the-shelf insights platform
DefineGlassInternal agent connected to data, research, strategy, codebase and design system. Acts as a PM's "tech lead" and builds prototypesPublic write-ups describe one Okta SSO login connecting 30+ toolsConnect an agent to your systems before asking it what to build
BuildInspectRamp's own coding agent. Works in Slack and returns a deploy preview75% of PRs; 1M+ sessions; ~1,000 PRs last month from non-engineersFollow Ramp's published spec (Modal sandboxes plus the open-source OpenCode framework)
BuildReview BuddyCode review agent that knows the codebase, runs quality and security checks, routes to the right human, and sees the generating prompts93% of PRs handled automaticallyGive a review agent codebase context and prompt provenance, plus a routing rule for human escalation
BuildTestoBrowser-based QA agent that runs the product in ~100 production-data combinations and returns bugs and design feedback425 bugs caught in 30 daysPair a browser-automation agent with production-derived test states
CoordinateGadgetAnswers questions directed at PMs using the formal record (Notion, Slack, Linear, tickets)85% of questions to PMs answered by AIMake the organization legible to agents, then route unanswerable questions to humans
ImproveAutonomous UX loopAI runs the loop for small issues: route, dedupe, rank, plan, human approval in Slack, code, test, ship, update docs60% of flagged UX issues fixed within 24 hoursChain intake, backlog matching and a coding agent behind a single Slack approval step

Note: every number comes from Geoff Charles's talk except the Glass architecture detail, which comes from public write-ups.

Step 1 — Identify: Ramp's Customer Insight Agent

Ramp's customer insight agent pulls customer pain from every company data source, clusters it around shared context, and tells the team which problems are real and which customers to talk to.

What it does

At most B2B companies, customer pain is scattered across Gong calls, Zendesk tickets, LogRocket session replays, surveys and angry emails sent straight to executives. Ramp's agent pulls from all of them. It uses traditional ETL and pipelines, vector search, and clustering around shared context. It also knows Ramp's product, teams and features, so feedback can be tied to the right feature and team instead of sitting in a generic pile.

The origin story

The first version was a Slack "hate channel" that posted customer quotes every day. It worked until it "got out of hand." The volume outgrew anyone's ability to read it, so the team rebuilt it from scratch as a proper system.

The number

According to the talk, a 1M-token context window covers less than 0.5% of Ramp's Gong transcripts. For a rough sense of scale (our estimate, not a Ramp figure): at about 150 spoken words per minute and 1.3 tokens per word, a 30-minute call is around 6,000 tokens. One million tokens holds roughly 170 calls. If that is under 0.5% of the corpus, Ramp is sitting on 200M+ tokens, or tens of thousands of calls. Pasting transcripts into a chatbot does not work at that size. Retrieval and clustering are required.

How it is surfaced

  • A Slack agent anyone can question
  • An HTML dashboard
  • A daily "hate podcast" of roughly 100 customers complaining

The goal is not to replace talking to customers. It is to know exactly which customers to talk to, because every insight traces back to its source.

How to copy it

  1. Inventory every place customer pain lives: calls, tickets, session replays, surveys and executive inboxes.
  2. Pipe all of them into one store tagged with product, team and feature metadata.
  3. Use vector search plus clustering instead of stuffing everything into one context window.
  4. Deliver insights where people already work: Slack, a dashboard or an audio digest.
  5. Keep every insight linked to the original quote and customer.

For the full build, see our step-by-step guide to building a customer insight agent like Ramp's.

If you would rather not maintain pipelines, BuildBetter is an off-the-shelf option for this step. It pulls calls, Slack threads, tickets and surveys together through 100+ integrations, analyzes each piece of feedback individually, and keeps every output traceable to the customer who said it. It is not something Ramp uses; it is a way to get the same capability without an in-house data team. We compare options in Ramp customer insight agent alternatives.

Step 2 — Define: Ramp Glass

Ramp Glass is an internal AI agent built on one premise: "What do you want to build?" is the wrong first prompt, so the first move is to connect the AI to your systems.

What it connects to

  • Snowflake, for quantitative data
  • User research, for qualitative evidence
  • Product strategy
  • Ramp's spec format
  • The codebase
  • The design system
  • Product principles

The "tech lead" role

With that context, a PM can ask Glass "is this possible?" or "would this break something?" and get an answer grounded in the actual code. Glass also builds working prototypes inside the real product, not in a disconnected mockup tool. That changes what a PM brings to engineering.

The next contract between PM and engineering

Charles described the handoff engineers actually want in three parts:

  1. Evidence: qual and quant proof that the problem is real.
  2. Requirements a coding agent can use: specific enough for Inspect to act on.
  3. A prototype for inspiration: something to react to, not a final build.

A long spec alone is not enough. Neither is a prototype alone. The value comes from all three together.

Architecture (from public write-ups, not the talk)

Public coverage describes Glass as built on Anthropic's Claude Agent SDK, with a single Okta SSO login connecting 30+ tools including Salesforce, Snowflake, Gong, Slack and Figma. The talk focused on what Glass does rather than its stack, so treat these details as secondary-source.

How to copy it

  1. Write down your spec format and product principles so an agent can follow them.
  2. Connect one quant source and one qual source first.
  3. Give the agent read access to the codebase and design system so its answers are grounded.
  4. Change your PRD template to the three-part contract.
  5. Require every spec to cite evidence.

More options are in Ramp Glass alternatives. For the evidence-and-requirements half of Define, BuildBetter generates PRDs and tickets grounded in real customer conversations, with each requirement linked to its source. To be clear about scope: building prototypes inside your own codebase is outside what it does.

Step 3 — Build: Inspect, Review Buddy and Testo

Once specs arrive faster, the bottleneck moves to writing, reviewing and testing code. Ramp built one agent for each of those jobs.

Ramp Inspect: what it does

Inspect is Ramp's own background coding agent. Ramp built it in-house to control the environment around the model: the sandbox, the tools, the codebase context and the verification loop. It runs inside Slack and returns a deploy preview, so the person who asked for a change can check the result in a browser.

Inspect: the numbers

  • 1M+ sessions
  • 75% of Ramp's pull requests are built by Inspect
  • ~1,000 PRs last month were submitted by non-engineers

The deploy preview explains that last figure. Non-engineers such as PMs verify behavior, not code, so they can ship small changes without learning to read a diff.

Inspect: the stack

According to Ramp's engineering blog, Inspect runs on Modal sandboxes with the open-source OpenCode agent framework. Ramp published the design in "Why we built our background agent" so other teams can replicate it. The model is swappable. The environment is where the advantage sits.

Inspect: how to copy it

  1. Read Ramp's published spec before designing anything.
  2. Run agents in isolated sandboxes, not on developer laptops.
  3. Start from an open-source agent framework rather than writing one from scratch.
  4. Trigger it from Slack so non-engineers can use it.
  5. Return a deploy preview so requesters can verify output without reading code.

Review Buddy: what it does and the number

When code generation gets cheap, review becomes the constraint. Review Buddy knows the codebase, runs quality and security checks, routes to the right human reviewer, and sees the prompts that produced the code. 93% of PRs are handled automatically, which lets senior engineers spend their time on the 7% that carry real risk.

Prompt provenance is the underrated part. A reviewer that can read the prompt behind a diff can judge whether the implementation matches the intent, not only whether the syntax is clean.

Review Buddy: how to copy the pattern

  1. Give the reviewer full codebase context, not just the diff.
  2. Attach the generating prompt to every agent-written PR.
  3. Encode your quality and security checks as explicit rules.
  4. Define routing rules for which human reviews which risk category.
  5. Measure the share auto-approved versus escalated.

Testo: what it does and the number

Testo is a browser-based QA agent. It spins the product up in about 100 combinations derived from production data, clicks through like a user, and returns both bugs and design feedback. It caught 425 bugs in the last 30 days.

Testo: how to copy the pattern

  1. Derive test states from real production data shapes.
  2. Let the agent navigate the UI like a user instead of relying only on scripted assertions.
  3. Ask for design feedback, not only pass/fail.
  4. Run it on every deploy preview the coding agent produces.

Step 4 — Coordinate: Ramp Gadget

Gadget is built on the idea that "every question is an API." When building gets faster, human attention becomes the bottleneck, and PMs spend their days answering the same questions over and over.

What it does

Gadget connects the intent behind a question to the formal record: roadmaps, specs and customer calls stored in Notion, Slack, Linear and tickets. It only works because Ramp writes things down. In Charles's words, "The organization needs to be legible to your agents."

What Gadget handles

  • Status questions, answered with evidence
  • Roadmap updates
  • Pinging owners who are late
  • Sales questions about availability, pricing and use case
  • Routing anything it cannot answer to the right person
  • Drafting help-center articles, blog posts and customer emails at launch

The number

85% of questions asked to PMs are answered by AI. The rest are answered and fed back into the system, so coverage keeps growing.

How to copy it

  1. Log the questions PMs get for two weeks and categorize them.
  2. Make sure the answers exist in a written system of record.
  3. Connect the agent to Notion, Slack, Linear and tickets, or your equivalents.
  4. Require every answer to cite its source.
  5. Build an explicit "route to a human" path for questions it cannot answer.

The two-week log usually shows that a handful of question types make up most of the volume. Start there.

Step 5 — Improve: The Autonomous UX Fix Loop

For small issues, Ramp lets AI run the entire loop from report to release, with one human approval in the middle.

What it does

When a UX issue comes in, the system:

  1. Routes it to the right team
  2. Matches it to the Linear backlog
  3. Dedupes and counts related reports
  4. Ranks it
  5. Writes a plan
  6. Asks for human approval in Slack ("I'm ready to code")
  7. Writes the code
  8. Runs tests and CI/CD
  9. Updates the knowledge base at launch

The number

60% of UX issues flagged by a customer, salesperson, CX or internal team are fixed within 24 hours.

This step chains the earlier ones. Intake comes from Identify. Context comes from Define. Coding and QA come from Build. Documentation comes from Coordinate. You cannot build this loop first, because it depends on everything else already working.

How to copy it

  1. Define "small issue" narrowly at first.
  2. Automate dedupe and backlog matching before automating code.
  3. Keep exactly one human approval gate in Slack.
  4. Require tests and CI to pass before merge.
  5. Close the loop by updating docs and telling the person who reported it.

That last step is the one most teams skip, and it is the one customers notice. Read the full breakdown in how Ramp fixes UX issues in 24 hours.

Where to Start If You Are Not Ramp

Constraints force you to choose where to compete. Charles told the story of Audi at Le Mans: instead of chasing every advantage, Audi picked one dimension, fuel efficiency, which meant fewer pit stops. That single choice decided races.

Rule 1: Pick one bottleneck

Do not try to build all seven tools at once. Ramp built these over time, each in response to a constraint it could measure.

Rule 2: Most teams should start at Identify

Every later step depends on knowing which problems are real. Faster code aimed at the wrong problem is waste, delivered sooner.

Find your current bottleneck

  • Where do requests wait longest?
  • Which role is always overloaded?
  • Which handoff causes the most rework?

Suggested sequence

Identify, then Define, then Coordinate, then Build, then Improve. If your data shows a different constraint, follow the data. Bottlenecks tend to migrate in a predictable order: faster specs expose coding capacity, faster coding exposes review, faster review exposes QA, and faster shipping exposes coordination.

Build vs. buy

Ramp built everything in-house. Smaller teams usually cannot staff that, and the Identify and Define layers are the easiest to buy.

OptionBest fitSteps coveredTrade-off
BuildBetter (off-the-shelf)B2B product teams that want calls, tickets, Slack and surveys unified without building pipelinesIdentify, plus evidence-backed PRDs and tickets for DefineDoes not prototype inside your codebase
Build in-house (Ramp's path)Teams with data engineering capacity and very specific needsAny stepOngoing maintenance and slower time to first value
Dedicated research repository or enterprise survey toolTeams whose main need is organizing research studies or running large surveysParts of IdentifyNarrower source coverage

BuildBetter is one reasonable starting point for the Identify layer. Be honest about your needs, though: if your work centers on research studies or large-scale surveys, a specialized tool may fit better.

"Everything you've seen here is outdated." — Geoff Charles

Design for the bottleneck to move again. Whatever you fix this quarter will expose the next constraint.

The Three PM Tracks in an AI Software Factory

Charles described three tracks for PMs as the factory takes over routine work.

  • Technical PM: builds and maintains the factory itself: the agents, the integrations and the scaffolding that lets them work on real systems.
  • Taste-maker: owns judgment on what is good. This role grows in value as output volume rises, because someone has to decide which of 100 prototypes deserves to ship.
  • GM: owns business outcomes across the whole loop, from customer pain to revenue.

Self-assessment

  • Do you enjoy wiring systems together and debugging agent behavior? You are likely a technical PM, and Build or Improve is your step.
  • Do people bring you work to ask "is this good?" You are likely a taste-maker, and Define is your step.
  • Do you think in terms of revenue, retention and churn? You are likely a GM, and Identify plus Coordinate are yours.

Pick one track and one factory step. That is where you should build depth over the next year.

FAQ: Ramp's AI Tools

Can you buy Ramp's AI tools like Inspect, Glass, Review Buddy, Testo or Gadget?

No. All of them are internal tools built by Ramp for its own teams and are not commercial products. You can replicate the patterns, and Ramp has published an engineering spec for Inspect.

Are any of Ramp's AI tools open source?

The tools themselves are not open source. Inspect is built on the open-source OpenCode agent framework running in Modal sandboxes, and Ramp documented the design in "Why we built our background agent" on builders.ramp.com so others can replicate it.

What percentage of Ramp's code is written by AI agents?

According to Geoff Charles's September 2026 talk, 75% of Ramp's pull requests are built by Inspect, and about 1,000 PRs in the prior month came from non-engineers.

What is Ramp Glass built on?

Public write-ups describe Glass as built on Anthropic's Claude Agent SDK, with a single Okta SSO login connecting 30+ internal tools such as Salesforce, Snowflake, Gong, Slack and Figma. The talk itself focused on what Glass does (connecting to data, research, strategy, codebase and design system) rather than its stack.

What is an AI software factory?

It is the system of agents and workflows that builds the product, as opposed to the product itself. In Ramp's version it has five steps (Identify, Define, Build, Coordinate, Improve), each with AI tooling aimed at the current bottleneck.

Where should a smaller team start?

At Identify. Build or buy a customer insight layer first, because every later step depends on knowing which problems are real. BuildBetter is one off-the-shelf option for this step; it has no relationship with Ramp.

Start Your Own Factory at Step One

Ramp's factory starts with knowing exactly what customers are struggling with. BuildBetter gives B2B product teams that layer without building it themselves. Calls, tickets, Slack threads and surveys come into one place, and every piece of feedback is analyzed individually for severity and business impact. From there you get PRDs and tickets with the evidence attached, and customers hear back automatically when the fix ships. It is SOC 2 Type II, HIPAA and GDPR compliant, and more than 30,000 teams use it.

Make churn optional. Book a demo.