Ramp's AI Product Factory: Every Internal Tool Explained (2026)
A tool-by-tool guide to Ramp's AI software factory from Geoff Charles's 2026 talk: Inspect, Glass, Review Buddy, Testo, Gadget, and how to copy each one.
In September 2026, Geoff Charles, Chief Product Officer at Ramp, gave a talk at the Lenny and Friends Summit called "The limiting factor — how to design an AI software factory for speed." This guide from the BuildBetter team goes through every internal tool he showed, what each one does, the metric he attached to it, and how you could build the same pattern yourself. His thesis fits in one sentence: AI does not remove the bottleneck, it moves it, and the teams that win are the ones that find the next bottleneck fastest.
That idea echoes Eliyahu Goldratt's Theory of Constraints. A system's throughput is set by its tightest constraint, and when you fix that constraint, a different one takes its place. Charles asked PMs to invest in the factory that builds the product, not only in the product itself. His closing line was blunt: "I want you to copy us." This page is meant to help you do that.
Two disclosures first. Every tool below is internal to Ramp. Ramp sells corporate cards and spend management software, not developer or PM tooling, so none of these tools is for sale. The pattern behind each one is what you can copy. Second, Charles gave his own caveat: "Everything you've seen here is outdated." Read the tools as a snapshot of a system that keeps changing, not as a fixed blueprint.
The structure: five steps (Identify, Define, Build, Coordinate, Improve), seven tools, one summary table, and a "how to copy it" list for each tool.
Every Ramp AI Tool at a Glance (Summary Table)
Ramp's AI software factory has seven tools spread across five steps. Each one exists to relieve the constraint created by speeding up the step before it.
| Step | Tool | What it does | Headline number | How to replicate it |
|---|---|---|---|---|
| Identify | Customer insight agent | Pulls customer pain from every data source, then clusters and ranks it | A 1M-token context window is less than 0.5% of Ramp's Gong transcripts | Build ETL, vector search and clustering over all feedback sources, or use an off-the-shelf insights platform |
| Define | Glass | Internal agent connected to data, research, strategy, codebase and design system. Acts as a PM's "tech lead" and builds prototypes | Public write-ups describe one Okta SSO login connecting 30+ tools | Connect an agent to your systems before asking it what to build |
| Build | Inspect | Ramp's own coding agent. Works in Slack and returns a deploy preview | 75% of PRs; 1M+ sessions; ~1,000 PRs last month from non-engineers | Follow Ramp's published spec (Modal sandboxes plus the open-source OpenCode framework) |
| Build | Review Buddy | Code review agent that knows the codebase, runs quality and security checks, routes to the right human, and sees the generating prompts | 93% of PRs handled automatically | Give a review agent codebase context and prompt provenance, plus a routing rule for human escalation |
| Build | Testo | Browser-based QA agent that runs the product in ~100 production-data combinations and returns bugs and design feedback | 425 bugs caught in 30 days | Pair a browser-automation agent with production-derived test states |
| Coordinate | Gadget | Answers questions directed at PMs using the formal record (Notion, Slack, Linear, tickets) | 85% of questions to PMs answered by AI | Make the organization legible to agents, then route unanswerable questions to humans |
| Improve | Autonomous UX loop | AI runs the loop for small issues: route, dedupe, rank, plan, human approval in Slack, code, test, ship, update docs | 60% of flagged UX issues fixed within 24 hours | Chain intake, backlog matching and a coding agent behind a single Slack approval step |
Note: every number comes from Geoff Charles's talk except the Glass architecture detail, which comes from public write-ups.
Step 1 — Identify: Ramp's Customer Insight Agent
Ramp's customer insight agent pulls customer pain from every company data source, clusters it around shared context, and tells the team which problems are real and which customers to talk to.
What it does
At most B2B companies, customer pain is scattered across Gong calls, Zendesk tickets, LogRocket session replays, surveys and angry emails sent straight to executives. Ramp's agent pulls from all of them. It uses traditional ETL and pipelines, vector search, and clustering around shared context. It also knows Ramp's product, teams and features, so feedback can be tied to the right feature and team instead of sitting in a generic pile.
The origin story
The first version was a Slack "hate channel" that posted customer quotes every day. It worked until it "got out of hand." The volume outgrew anyone's ability to read it, so the team rebuilt it from scratch as a proper system.
The number
According to the talk, a 1M-token context window covers less than 0.5% of Ramp's Gong transcripts. For a rough sense of scale (our estimate, not a Ramp figure): at about 150 spoken words per minute and 1.3 tokens per word, a 30-minute call is around 6,000 tokens. One million tokens holds roughly 170 calls. If that is under 0.5% of the corpus, Ramp is sitting on 200M+ tokens, or tens of thousands of calls. Pasting transcripts into a chatbot does not work at that size. Retrieval and clustering are required.
How it is surfaced
- A Slack agent anyone can question
- An HTML dashboard
- A daily "hate podcast" of roughly 100 customers complaining
The goal is not to replace talking to customers. It is to know exactly which customers to talk to, because every insight traces back to its source.
How to copy it
- Inventory every place customer pain lives: calls, tickets, session replays, surveys and executive inboxes.
- Pipe all of them into one store tagged with product, team and feature metadata.
- Use vector search plus clustering instead of stuffing everything into one context window.
- Deliver insights where people already work: Slack, a dashboard or an audio digest.
- Keep every insight linked to the original quote and customer.
For the full build, see our step-by-step guide to building a customer insight agent like Ramp's.
If you would rather not maintain pipelines, BuildBetter is an off-the-shelf option for this step. It pulls calls, Slack threads, tickets and surveys together through 100+ integrations, analyzes each piece of feedback individually, and keeps every output traceable to the customer who said it. It is not something Ramp uses; it is a way to get the same capability without an in-house data team. We compare options in Ramp customer insight agent alternatives.
Step 2 — Define: Ramp Glass
Ramp Glass is an internal AI agent built on one premise: "What do you want to build?" is the wrong first prompt, so the first move is to connect the AI to your systems.
What it connects to
- Snowflake, for quantitative data
- User research, for qualitative evidence
- Product strategy
- Ramp's spec format
- The codebase
- The design system
- Product principles
The "tech lead" role
With that context, a PM can ask Glass "is this possible?" or "would this break something?" and get an answer grounded in the actual code. Glass also builds working prototypes inside the real product, not in a disconnected mockup tool. That changes what a PM brings to engineering.
The next contract between PM and engineering
Charles described the handoff engineers actually want in three parts:
- Evidence: qual and quant proof that the problem is real.
- Requirements a coding agent can use: specific enough for Inspect to act on.
- A prototype for inspiration: something to react to, not a final build.
A long spec alone is not enough. Neither is a prototype alone. The value comes from all three together.
Architecture (from public write-ups, not the talk)
Public coverage describes Glass as built on Anthropic's Claude Agent SDK, with a single Okta SSO login connecting 30+ tools including Salesforce, Snowflake, Gong, Slack and Figma. The talk focused on what Glass does rather than its stack, so treat these details as secondary-source.
How to copy it
- Write down your spec format and product principles so an agent can follow them.
- Connect one quant source and one qual source first.
- Give the agent read access to the codebase and design system so its answers are grounded.
- Change your PRD template to the three-part contract.
- Require every spec to cite evidence.
More options are in Ramp Glass alternatives. For the evidence-and-requirements half of Define, BuildBetter generates PRDs and tickets grounded in real customer conversations, with each requirement linked to its source. To be clear about scope: building prototypes inside your own codebase is outside what it does.
Step 3 — Build: Inspect, Review Buddy and Testo
Once specs arrive faster, the bottleneck moves to writing, reviewing and testing code. Ramp built one agent for each of those jobs.
Ramp Inspect: what it does
Inspect is Ramp's own background coding agent. Ramp built it in-house to control the environment around the model: the sandbox, the tools, the codebase context and the verification loop. It runs inside Slack and returns a deploy preview, so the person who asked for a change can check the result in a browser.
Inspect: the numbers
- 1M+ sessions
- 75% of Ramp's pull requests are built by Inspect
- ~1,000 PRs last month were submitted by non-engineers
The deploy preview explains that last figure. Non-engineers such as PMs verify behavior, not code, so they can ship small changes without learning to read a diff.
Inspect: the stack
According to Ramp's engineering blog, Inspect runs on Modal sandboxes with the open-source OpenCode agent framework. Ramp published the design in "Why we built our background agent" so other teams can replicate it. The model is swappable. The environment is where the advantage sits.
Inspect: how to copy it
- Read Ramp's published spec before designing anything.
- Run agents in isolated sandboxes, not on developer laptops.
- Start from an open-source agent framework rather than writing one from scratch.
- Trigger it from Slack so non-engineers can use it.
- Return a deploy preview so requesters can verify output without reading code.
Review Buddy: what it does and the number
When code generation gets cheap, review becomes the constraint. Review Buddy knows the codebase, runs quality and security checks, routes to the right human reviewer, and sees the prompts that produced the code. 93% of PRs are handled automatically, which lets senior engineers spend their time on the 7% that carry real risk.
Prompt provenance is the underrated part. A reviewer that can read the prompt behind a diff can judge whether the implementation matches the intent, not only whether the syntax is clean.
Review Buddy: how to copy the pattern
- Give the reviewer full codebase context, not just the diff.
- Attach the generating prompt to every agent-written PR.
- Encode your quality and security checks as explicit rules.
- Define routing rules for which human reviews which risk category.
- Measure the share auto-approved versus escalated.
Testo: what it does and the number
Testo is a browser-based QA agent. It spins the product up in about 100 combinations derived from production data, clicks through like a user, and returns both bugs and design feedback. It caught 425 bugs in the last 30 days.
Testo: how to copy the pattern
- Derive test states from real production data shapes.
- Let the agent navigate the UI like a user instead of relying only on scripted assertions.
- Ask for design feedback, not only pass/fail.
- Run it on every deploy preview the coding agent produces.
Step 4 — Coordinate: Ramp Gadget
Gadget is built on the idea that "every question is an API." When building gets faster, human attention becomes the bottleneck, and PMs spend their days answering the same questions over and over.
What it does
Gadget connects the intent behind a question to the formal record: roadmaps, specs and customer calls stored in Notion, Slack, Linear and tickets. It only works because Ramp writes things down. In Charles's words, "The organization needs to be legible to your agents."
What Gadget handles
- Status questions, answered with evidence
- Roadmap updates
- Pinging owners who are late
- Sales questions about availability, pricing and use case
- Routing anything it cannot answer to the right person
- Drafting help-center articles, blog posts and customer emails at launch
The number
85% of questions asked to PMs are answered by AI. The rest are answered and fed back into the system, so coverage keeps growing.
How to copy it
- Log the questions PMs get for two weeks and categorize them.
- Make sure the answers exist in a written system of record.
- Connect the agent to Notion, Slack, Linear and tickets, or your equivalents.
- Require every answer to cite its source.
- Build an explicit "route to a human" path for questions it cannot answer.
The two-week log usually shows that a handful of question types make up most of the volume. Start there.
Step 5 — Improve: The Autonomous UX Fix Loop
For small issues, Ramp lets AI run the entire loop from report to release, with one human approval in the middle.
What it does
When a UX issue comes in, the system:
- Routes it to the right team
- Matches it to the Linear backlog
- Dedupes and counts related reports
- Ranks it
- Writes a plan
- Asks for human approval in Slack ("I'm ready to code")
- Writes the code
- Runs tests and CI/CD
- Updates the knowledge base at launch
The number
60% of UX issues flagged by a customer, salesperson, CX or internal team are fixed within 24 hours.
This step chains the earlier ones. Intake comes from Identify. Context comes from Define. Coding and QA come from Build. Documentation comes from Coordinate. You cannot build this loop first, because it depends on everything else already working.
How to copy it
- Define "small issue" narrowly at first.
- Automate dedupe and backlog matching before automating code.
- Keep exactly one human approval gate in Slack.
- Require tests and CI to pass before merge.
- Close the loop by updating docs and telling the person who reported it.
That last step is the one most teams skip, and it is the one customers notice. Read the full breakdown in how Ramp fixes UX issues in 24 hours.
Where to Start If You Are Not Ramp
Constraints force you to choose where to compete. Charles told the story of Audi at Le Mans: instead of chasing every advantage, Audi picked one dimension, fuel efficiency, which meant fewer pit stops. That single choice decided races.
Rule 1: Pick one bottleneck
Do not try to build all seven tools at once. Ramp built these over time, each in response to a constraint it could measure.
Rule 2: Most teams should start at Identify
Every later step depends on knowing which problems are real. Faster code aimed at the wrong problem is waste, delivered sooner.
Find your current bottleneck
- Where do requests wait longest?
- Which role is always overloaded?
- Which handoff causes the most rework?
Suggested sequence
Identify, then Define, then Coordinate, then Build, then Improve. If your data shows a different constraint, follow the data. Bottlenecks tend to migrate in a predictable order: faster specs expose coding capacity, faster coding exposes review, faster review exposes QA, and faster shipping exposes coordination.
Build vs. buy
Ramp built everything in-house. Smaller teams usually cannot staff that, and the Identify and Define layers are the easiest to buy.
| Option | Best fit | Steps covered | Trade-off |
|---|---|---|---|
| BuildBetter (off-the-shelf) | B2B product teams that want calls, tickets, Slack and surveys unified without building pipelines | Identify, plus evidence-backed PRDs and tickets for Define | Does not prototype inside your codebase |
| Build in-house (Ramp's path) | Teams with data engineering capacity and very specific needs | Any step | Ongoing maintenance and slower time to first value |
| Dedicated research repository or enterprise survey tool | Teams whose main need is organizing research studies or running large surveys | Parts of Identify | Narrower source coverage |
BuildBetter is one reasonable starting point for the Identify layer. Be honest about your needs, though: if your work centers on research studies or large-scale surveys, a specialized tool may fit better.
"Everything you've seen here is outdated." — Geoff Charles
Design for the bottleneck to move again. Whatever you fix this quarter will expose the next constraint.
The Three PM Tracks in an AI Software Factory
Charles described three tracks for PMs as the factory takes over routine work.
- Technical PM: builds and maintains the factory itself: the agents, the integrations and the scaffolding that lets them work on real systems.
- Taste-maker: owns judgment on what is good. This role grows in value as output volume rises, because someone has to decide which of 100 prototypes deserves to ship.
- GM: owns business outcomes across the whole loop, from customer pain to revenue.
Self-assessment
- Do you enjoy wiring systems together and debugging agent behavior? You are likely a technical PM, and Build or Improve is your step.
- Do people bring you work to ask "is this good?" You are likely a taste-maker, and Define is your step.
- Do you think in terms of revenue, retention and churn? You are likely a GM, and Identify plus Coordinate are yours.
Pick one track and one factory step. That is where you should build depth over the next year.
FAQ: Ramp's AI Tools
Can you buy Ramp's AI tools like Inspect, Glass, Review Buddy, Testo or Gadget?
No. All of them are internal tools built by Ramp for its own teams and are not commercial products. You can replicate the patterns, and Ramp has published an engineering spec for Inspect.
Are any of Ramp's AI tools open source?
The tools themselves are not open source. Inspect is built on the open-source OpenCode agent framework running in Modal sandboxes, and Ramp documented the design in "Why we built our background agent" on builders.ramp.com so others can replicate it.
What percentage of Ramp's code is written by AI agents?
According to Geoff Charles's September 2026 talk, 75% of Ramp's pull requests are built by Inspect, and about 1,000 PRs in the prior month came from non-engineers.
What is Ramp Glass built on?
Public write-ups describe Glass as built on Anthropic's Claude Agent SDK, with a single Okta SSO login connecting 30+ internal tools such as Salesforce, Snowflake, Gong, Slack and Figma. The talk itself focused on what Glass does (connecting to data, research, strategy, codebase and design system) rather than its stack.
What is an AI software factory?
It is the system of agents and workflows that builds the product, as opposed to the product itself. In Ramp's version it has five steps (Identify, Define, Build, Coordinate, Improve), each with AI tooling aimed at the current bottleneck.
Where should a smaller team start?
At Identify. Build or buy a customer insight layer first, because every later step depends on knowing which problems are real. BuildBetter is one off-the-shelf option for this step; it has no relationship with Ramp.
Start Your Own Factory at Step One
Ramp's factory starts with knowing exactly what customers are struggling with. BuildBetter gives B2B product teams that layer without building it themselves. Calls, tickets, Slack threads and surveys come into one place, and every piece of feedback is analyzed individually for severity and business impact. From there you get PRDs and tickets with the evidence attached, and customers hear back automatically when the fix ships. It is SOC 2 Type II, HIPAA and GDPR compliant, and more than 30,000 teams use it.