How to Build a Customer Insight Agent Like Ramp: 2026 Guide

A six-stage architecture for a Ramp-style customer insight agent: connectors, pgvector, BERTopic clustering, cited answers, and an honest build vs buy.

How to Build a Customer Insight Agent Like Ramp: 2026 Guide

At the Lenny and Friends Summit in September 2026, Geoff Charles, Chief Product Officer at Ramp, gave a talk titled "The limiting factor — how to design an AI software factory for speed." He ended it with a direct request: "I want you to copy us." This guide takes him at his word. We're BuildBetter, and we build customer insight software for B2B product teams. That means we've spent years on the problems you'll hit if you build this yourself. Below is a reference architecture you can build, the open-source defaults we'd pick, the failure modes that don't show up in demos, and a build-vs-buy threshold.

Ramp's version started small. It began as a Slack "hate channel" that posted customer quotes every day. Charles said it "got out of hand," so the team rebuilt it from scratch. The new version is an agent that pulls from all of Ramp's data sources using ETL pipelines, vector search and clustering. It knows the product, the teams and the features.

The agent does not replace talking to customers. It tells you exactly which customers to talk to, because every insight traces back to its source. Everything else in this guide supports that goal.

Ramp's agent is an internal tool. It is not for sale, and no vendor sells or licenses it, which is why this guide exists. The insight agent is step 1 (IDENTIFY) of the five-stage factory Charles described. For the full system, see our breakdown of Ramp's AI product factory and its internal tools.

What Is a Customer Insight Agent?

A customer insight agent is a system that continuously ingests customer conversations and feedback from every source, groups them into themes mapped to your product areas, and answers questions with citations back to the original quote, account and timestamp.

It differs from a feedback inbox or a shared spreadsheet of quotes in three ways:

  • Continuous, not batch. New calls and tickets flow in daily. Nobody runs a quarterly export and a synthesis sprint.
  • Structured and ranked, not a raw stream. Feedback is deduplicated, attributed to accounts, clustered into themes and ranked by impact. It is never a firehose of quotes.
  • Queryable and traceable, not a summary you have to trust. Every claim links to the verbatim, the account and the moment it was said.

Typical inputs are the same source types Ramp named:

  • Sales and customer success calls (Gong, Zoom recordings)
  • Support tickets (Zendesk, Intercom)
  • Session replay notes and flagged sessions (LogRocket)
  • Survey responses and open-text fields
  • Angry emails forwarded to executives

Typical outputs:

  • Ranked themes by product area and owning team
  • A list of affected accounts with revenue context (ARR, plan, renewal date)
  • Cited answers to ad-hoc questions, such as "What are enterprise accounts saying about approvals this month?"

The term overlaps with voice of customer AI and VoC analytics, but the emphasis is different. A customer insight agent treats unstructured conversation as the primary signal. Surveys are one input among several. It also treats traceability as a hard requirement.

Why "Just Paste It Into an LLM" Fails

At real feedback volume, the problem is retrieval and aggregation, not context length. Charles made this concrete: at Ramp, a 1M-token context window holds "less than 0.5% of Gong transcripts." Bigger models won't fix that. The corpus grows faster than the window.

To make the ratio tangible, here is our own estimate. Conversational speech runs about 150 words per minute, and English averages about 1.3 tokens per word. That puts one call-hour at roughly 12,000 tokens, so 1M tokens is about 80 hours of transcribed conversation. If 1M tokens is under 0.5% of Ramp's corpus, the corpus exceeds roughly 200M tokens, which is on the order of 15,000+ hours of recorded calls. A team doing 20 hours of calls a week outgrows a 1M-token window in about a month.

Even when your data fits, long-context prompting fails in three specific ways:

  1. It undercounts mid-context material. The "Lost in the Middle" study (Liu et al., Stanford, TACL 2024) found that LLM accuracy drops when relevant information sits in the middle of a long context rather than at the start or end. If the model quietly skips a third of your complaints, your frequency ranking is fiction.
  2. It gives no stable citations. A summary that says "several customers are frustrated with sync" doesn't tell you who to call.
  3. It forgets between sessions. There is no persistent store, so you can't ask whether a theme is growing.

What works instead is the stack Ramp described: ETL, vector search and clustering around shared product context. You need to find the relevant 0.5% and count patterns across the other 99.5%. The rest of this guide breaks that stack into six stages you can build.

The Reference Architecture: 6 Stages You Can Build

A working customer insight agent has six stages: connectors, a normalized schema, embeddings, clustering, a cited query layer and distribution surfaces. Each stage below follows the same format: what to build, the open-source default, and what quietly breaks.

  1. Sources and connectors — pull calls, tickets, replays, surveys and escalation emails
  2. ETL and a normalized record schema — one shape for every piece of feedback
  3. Chunking, embeddings and vector store — make feedback searchable by meaning
  4. Clustering and taxonomy mapping — turn chunks into named themes owned by teams
  5. Query layer with citations — answer questions with links to evidence
  6. Distribution surfaces — Slack, dashboard, daily audio
[Gong / Zoom]  [Zendesk / Intercom]  [LogRocket]  [Surveys]  [Exec inbox]  [CRM]
       \              |                  |           |           /          |
        +-------------+---- Stage 1: connectors -----+----------+           |
                              |                                            |
                 Stage 2: ETL + normalized records  <--- account + ARR join -+
                              |
                 Stage 3: chunks + embeddings (Postgres + pgvector)
                              |
                 Stage 4: clusters -> themes -> taxonomy (area / feature / team)
                              |
                 Stage 5: query layer (hybrid search + SQL counts + citations)
                              |
            Stage 6: Slack agent  |  HTML dashboard  |  daily audio digest

Disclaimer: this is a reference design inspired by the capabilities Ramp described publicly. It is not Ramp's actual implementation. The specific tools named are our recommendations, not Ramp's.

Stage 1: Sources and Connectors

The first stage pulls every customer-facing source into one place on a schedule, and it silently loses data unless you monitor record counts.

What to build

Set up scheduled pulls or webhooks from:

  • Your call recorder (Gong API or Zoom cloud recordings)
  • Your help desk (Zendesk, Intercom)
  • Session replay (LogRocket)
  • Survey tools
  • A forwarding inbox where executives send escalation emails

Add your CRM (Salesforce or HubSpot) as the join source for accounts and ARR. Start with the two highest-volume sources, which are usually calls and tickets. Add a third only after the themes from the first two look right to the PMs who own those areas.

Open-source default

Use dlt or Airbyte for connectors. For raw audio you record yourself, transcribe with Whisper or WhisperX. WhisperX adds word-level timestamps and diarization, which you'll need in Stage 2.

What quietly breaks

Connectors rarely fail loudly. API rate limits and pagination gaps drop records without raising an error. OAuth tokens expire over a holiday weekend. The fix is a per-source daily record-count check. Compare today's count to the trailing 14-day average and alert when it drops by more than a set threshold, such as 40%. Without this check, a theme can look like it's fading when the ticket connector stopped paging past 1,000 results.

Stage 2: ETL and the Normalized Record Schema

The normalized record schema is the most important design decision in the system. Every piece of feedback, whatever its source, becomes the same shape in one table.

What to build

FieldTypeExample
record_iduuid7f3c…
sourceenumgong / zendesk / logrocket / survey / email
source_urltextDeep link to the call timestamp or ticket
account_idtextCRM account ID
account_nametextAcme Logistics
arrnumeric84000
speaker_roleenumcustomer / rep / agent / unknown
speaker_nametextDana R. (ops lead)
occurred_attimestamptz2026-09-14T15:32:10Z
verbatimtext"We export to CSV every Friday because the sync breaks."
product_areatext (nullable, filled in Stage 4)Integrations > ERP sync
content_hashtextUsed for dedupe

Open-source default

Use Postgres as the single store and dbt for transforms. Keep raw landing tables per source, then build the normalized table in dbt so you can rebuild it when the schema changes.

What quietly breaks

  • Speaker attribution. If rep speech is labeled as customer speech, your agent will "discover" complaints your own sales team voiced. Gong labels speakers, but you still need to classify each one as internal or external. Mark any participant whose email domain matches your company domain as rep. Treat everyone else as a customer candidate, then map them to CRM contacts.
  • Account matching. Ticket emails and call attendees often fail to join to CRM accounts. Match by exact email first. Then match email domain to account domain, excluding gmail.com and other free domains. Then fall back to a manual override table. Track the unmatched rate as a metric.
  • PII. Redact emails, phone numbers and card data at ingest, before embedding, using Microsoft Presidio. Keep a reversible mapping only if you have a lawful need for it.

Stage 3: Chunking, Embeddings and the Vector Store

Chunk customer turns, not whole calls, and keep the vectors next to your metadata. A 45-minute call covers 5 to 10 topics. Embedding it whole produces a blurry average vector that matches everything weakly.

What to build

  • For calls, merge adjacent short customer turns until each chunk is about 100 to 400 tokens. Keep one preceding rep turn as context, so an answer to "what's your biggest pain?" stays interpretable.
  • For tickets, use one chunk per customer message.
  • Before embedding, prepend context such as [Customer: ops lead at Acme Logistics, $84k ARR, call about ERP sync]. Anthropic's Contextual Retrieval results show this substantially reduces retrieval failures. It also helps clustering separate similar-sounding complaints from different segments.
  • Store each vector alongside its Stage 2 record.

Open-source default

Use pgvector inside the same Postgres. It supports exact search plus approximate HNSW and IVFFlat indexes, with cosine, L2 and inner-product distance. The standard vector type indexes up to 2,000 dimensions, and halfvec goes up to 4,000. Pair it with an open embedding model such as BGE-M3 or nomic-embed-text-v1.5, or a hosted embedding API.

Keep vectors in the same database as your metadata until you have a measured reason not to. Most insight queries filter by account, date or ARR before they ever rank by similarity. That is one SQL query in Postgres, and a painful sync job across two systems.

What quietly breaks

  • Duplicates. The same complaint arrives as a ticket, comes up on a call, and gets forwarded by an exec. Use exact content hashes plus near-duplicate detection, such as a cosine similarity above 0.92 within the same account and a 7-day window. Otherwise one angry customer looks like five.
  • Re-embedding. Changing the model means re-embedding everything. Version your embeddings column (embedding_v2, model_name) so you can migrate without downtime.

Stage 4: Clustering Into Themes and Mapping to Product Areas

Clustering finds themes and taxonomy mapping gives each theme an owner. This is the layer that lets an agent understand the product, the teams and the features, which Charles said Ramp's agent does.

What to build

Run an unsupervised clustering pass to discover themes. Then run an LLM step that names each cluster and assigns it to your taxonomy of product areas, features and owning teams. Rank themes with an explicit, inspectable formula, for example:

score = unique_accounts × log(1 + total_ARR / 10,000) × exp(−days_since_last_mention / 14)

Show the formula in the dashboard. PMs should be able to argue with it, and they will.

Open-source default

BERTopic is the standard choice for feedback theme clustering. Its default pipeline runs sentence-transformer embeddings, then UMAP dimensionality reduction, then HDBSCAN clustering, then c-TF-IDF topic representations. It supports LLM-based labeling, seeded topics, zero-shot assignment to predefined topics and online updates. Constrain the LLM labeling pass to a taxonomy file, such as a YAML of areas → features → team owner, so it can't invent product areas.

What quietly breaks

  • Stale taxonomy. When you ship, rename or reorganize features, clusters get mapped to dead areas or dumped into "Other." Review the taxonomy monthly and alert when "Other" exceeds a set share, such as 15%.
  • Cluster drift. Re-run clustering on a schedule, and match new clusters to old ones by centroid similarity so theme IDs stay stable. Without stable IDs, every re-run resets your trend lines.

This taxonomy layer is where teams spend the most ongoing effort.

Stage 5: The Query Layer With Citations

The query layer follows one rule: no citation, no claim. The agent must answer "insufficient evidence" rather than generalize without source records.

What to build

Build an agent that answers natural-language questions using retrieval-augmented generation (RAG) with four parts:

  • Hybrid retrieval. Combine keyword search and vector search, merged with reciprocal rank fusion. Keywords catch exact feature names and error codes that embeddings miss.
  • Metadata filters. Exposed as tool calls: date range, ARR band, product area, segment.
  • A cross-encoder reranker. Reorders the top 50 to 100 candidates.
  • Cited answers. Each answer cites record IDs, rendered as links to the source.

Open-source default

Use Postgres full-text search plus pgvector for hybrid retrieval, an open cross-encoder reranker such as bge-reranker, and any capable LLM with tool calling.

Separate counting from retrieval

LLMs are good at naming, summarizing and answering. They are unreliable at counting across hundreds of passages. Every "how many accounts" or "is this growing" number should come from SQL over the clustered table, and the LLM should only narrate the result.

What quietly breaks

Citation drift is the failure to watch for. The answer cites a real record that does not support the claim. Add an automatic check that each cited quote actually contains the claimed content, using string match first and a small LLM judge second. Also sample 20 answers a week for human review.

Stage 6: Distribution Surfaces — Slack, Dashboard, Daily Audio

Insight that isn't pushed to people doesn't get used. Charles described three surfaces: a Slack agent you can ask questions, an HTML dashboard, and a daily "hate podcast" of roughly 100 customers complaining.

Slack agent

Build it with Slack Bolt. It should answer in-thread with citations and post weekly theme digests to each owning team's channel. Scope which channels it can read and write, and log every query. The query log tells you what people actually want to know.

Dashboard

Show ranked themes by product area and team, trend lines, and drill-down to verbatims and accounts. The open-source default is Metabase or Streamlit. A static HTML page regenerated nightly also works and is hard to break.

Daily audio digest

Select the top verbatims by rank, generate a short script, and render it with a text-to-speech model. Keep it brief and use real quotes, after PII redaction, so listeners hear the customer's own words. Hearing a customer say "we export to CSV every Friday" lands differently than reading a chart.

What quietly breaks

Attention decays. A raw firehose of quotes stops being read within weeks. Push ranked, deduplicated, owner-routed items instead of everything. If a team gets more than about 10 items a week, raise its threshold.

Worked Example: From 40 Raw Mentions to 5 Accounts to Call This Week

This is an illustrative, hypothetical scenario. No real company data is used.

Raw input: 40 mentions of ERP sync failures over 30 days: 22 call snippets, 14 Zendesk tickets and 4 exec-forwarded emails.

  1. Dedupe. Near-duplicate detection collapses cross-posted items. That leaves 31 unique mentions from 18 accounts.
  2. Speaker filter. Three snippets were a rep describing the issue, not a customer, so they are removed. That leaves 28 mentions from 17 accounts.
  3. Cluster and map. The theme is named "ERP sync fails silently; users fall back to CSV." It maps to Integrations > ERP sync, owned by the Integrations team.
  4. Rank. The theme scores high on unique accounts × ARR × recency and moves into the top three themes for the week.
  5. Output. The Integrations channel receives one ranked item:
    • A one-line summary
    • Three representative verbatims with links
    • The trend against the prior month
    • The five accounts to call this week: the highest-ARR accounts with the most recent mentions, each with a contact name and a timestamped link to the source

This is the point Ramp made: the agent does not replace the conversation. It tells you which conversations to have.

Build vs Buy: An Honest Threshold

Building a production customer insight agent takes roughly 8 to 16 engineer-weeks plus a quarter to half of an FTE to maintain. The estimates below are rough and vary with team and stack.

OptionScopeRough effort
Buy: BuildBetterCapture, unified schema, clustering, taxonomy, cited Q&A, Slack, plus PRDs, tickets and loop-closureConnect integrations; vendor maintains connectors, models and compliance
Build: MVP2 sources, clustering, Slack Q&A~3–6 engineer-weeks
Build: Production4–6 sources, PII redaction, taxonomy, dashboard, audio digest, evals~8–16 engineer-weeks
Build: Ongoing maintenanceConnectors, taxonomy, model upgrades, evals~0.25–0.5 FTE

Build it if

  • You have a platform or data engineer who can own it long term.
  • You already run Postgres and dbt.
  • Your sources are custom or unusual.
  • You need deep joins to proprietary product data.
  • You treat the system as a strategic capability, which is Ramp's "invest in the factory" argument.

Buy it if

  • You have fewer than about two engineers to spare.
  • You mostly use standard tools: Gong or Zoom, Zendesk or Intercom, Slack, Salesforce or HubSpot.
  • You need value in weeks rather than a quarter.
  • You need SOC 2 or HIPAA compliance without building it yourself.

Hidden costs teams underestimate

Four costs routinely take longer than the pipeline itself: speaker attribution quality, PII compliance review, taxonomy upkeep and answer-quality evals. None of them are one-time work.

If You'd Rather Not Build It: BuildBetter

BuildBetter covers the core of this pipeline out of the box and adds the step most internal tools skip: turning insight into shipped work.

  • Stage 1, capture. Records calls with or without a bot. It also pulls in tickets, Slack threads and surveys through 100+ integrations, including Zoom, Slack, Zendesk, Intercom, HubSpot, Salesforce and Jira.
  • Stages 2–4, unified schema and clustering. Signals analyzes each piece of feedback individually for severity, sentiment and business impact. Clusters & Insights groups signals into themes mapped to your taxonomy.
  • Stage 5, cited Q&A. Chat answers questions across every call, ticket and Slack thread, with citations back to the source.
  • Past dashboards. It produces PRDs, Linear and Jira tickets with full context, and loop-closure emails to the customers who asked.

It also combines external customer feedback with internal team activity, such as Slack discussions and call notes, in one place. Most homegrown agents only cover the external half. Compliance is built in: SOC 2 Type II, HIPAA and GDPR.

Some cases fit other tools better. Teams whose main job is a dedicated research repository, enterprise survey distribution at scale, or mining very large volumes of public reviews may prefer specialized tools. Ramp's agent itself is internal, and no vendor sells it. For a side-by-side look, see our guide to Ramp customer insight agent alternatives.

What Comes After Insight

The insight agent is the first bottleneck to fix, because every later stage is only as good as the evidence it starts with. In Ramp's factory, insight feeds four more stages:

  • DEFINE: Glass
  • BUILD: Inspect, Review Buddy and Testo
  • COORDINATE: Gadget
  • IMPROVE: an autonomous small-issue loop that, per Charles, fixes 60% of UX issues within 24 hours

Ramp published a spec for its background coding agent on its engineering blog: Why we built our background agent. For the IMPROVE loop, read how Ramp fixes UX issues in 24 hours.

FAQ

Is a customer insight agent the same as a VoC tool?

They overlap, but the emphasis differs. Traditional VoC tools center on surveys and scores (NPS, CSAT). An insight agent unifies unstructured conversations (calls, tickets, email, session replays), maps them to product areas and owning teams, and answers ad-hoc questions with traceable citations. Some platforms, such as BuildBetter, combine both.

How much data do I need before it's worth building?

This is a rule of thumb, not a statistic. Clustering and retrieval start paying off once no single person can read all incoming calls and tickets each week, or once feedback lives in three or more tools. About 20 hours of calls per week exceeds a 1M-token context window within roughly a month.

Can I do this with Claude or ChatGPT alone?

For a one-off analysis of a small sample, yes. For ongoing use, no. At Ramp, a 1M-token window held less than 0.5% of Gong transcripts. Long contexts also degrade on mid-context material, don't give stable citations, can't count reliably, and have no memory across weeks. LLMs are components of the system, not the system.

How do I keep PII out?

Redact at ingest, before embedding, with Presidio or a similar tool. Exclude payment and identity fields, restrict source links by role, and set retention windows. Don't send raw PII to third-party model APIs without a data processing agreement (DPA).

How long until it's useful?

An MVP on two sources can surface its first ranked themes within the 3–6 engineer-week estimate above. Trust takes longer. It comes from several weeks of citation-checked answers that PMs have spot-verified themselves.

Does it replace customer interviews?

No. Ramp was explicit that the goal is knowing which customers to talk to, not talking to them less.

Make churn optional.

You can spend a quarter building the six stages above, or you can have cited themes, ranked accounts and ready-to-ship tickets in front of your team this month. BuildBetter captures calls and pulls in tickets, Slack threads and surveys, and tells you which customers to call and what to build next.

Make churn optional. Book a demo →