A customer feedback system should connect each product request to its underlying customer evidence, show where that evidence came from, and remain affordable and repeatable to run. Brandon Galang describes building that system at Vercel and explains why evaluation and retrieval design matter after the first agent demo works.
Video not loading in LinkedIn or your in-app browser? Watch it on YouTube.
How does Brandon connect feedback to product decisions?
His system groups signals from customer calls, support cases, and field notes around product feature requests. The output keeps the underlying evidence beside each proposed match. A product manager can inspect why a request was grouped with a feature rather than accepting a summary or count on trust.
The interface shown in the talk uses synthetic examples. The design problem is real: teams often have feedback in several systems, and a single customer request can arrive through more than one channel.
What changes when the agent has to run repeatedly?
A prototype can retrieve a large candidate set and still produce a convincing answer. Brandon describes reducing the number of candidates and checking which useful evidence survives. The question is whether a cheaper run still finds the requests people care about. The numeric examples in the talk are working examples, not a benchmark for another team.
This gives the system two evaluation loops. One checks the quality of feedback-to-feature matches. The other checks whether retrieval costs and latency make the process practical to run on a schedule. Both need the original source and a way to inspect a wrong match.
Who decides whether a match is right?
Product, sales, and support may mean different things by the same feature name. Brandon's Q&A makes that disagreement visible. Build a small review set with examples each team accepts, and keep disputed cases instead of forcing one label into the report. A model score is only as useful as the decision it is meant to support.
Start with one feature area. Record the source, account context, proposed match, and reviewer decision. Then change retrieval or labeling based on actual misses, not on how plausible the generated summary sounds.
Watch and read more
Watch the speaker recording. Deepline's GTM data infrastructure guide covers the data foundations for source-linked workflows.
Frequently asked questions
What should a customer feedback system retain?
Keep the original signal, source, account context, proposed feature match, and the reviewer's decision.
Why evaluate an agent-built feedback system?
Evaluation shows whether the matches help the teams reading them and whether the retrieval process can run at a practical cost.
Full working transcript
The transcript is collapsed by default. Expand it to read the talk. Private names are omitted where needed.
Show full transcript
This is a working transcript with chapter timestamps. It may contain transcription errors. Private names have been removed; check the video before quoting a precise phrase.
00:00:00 The customer feedback system
00:00:00 Hello everybody. Okay. Okay. So my presentation is titled Your Agent Works, Now Make It a System. I'm going to be talking to you all about a customer feedback system I built for Vercel. Pretty new to the company, to be honest, a few months in. And I wanted to share a bit about my starting point.
00:00:36 So go-to-market engineering is pretty new for me, go-to-market as well as general. I come from product management and like in the past year and a half made the transition into being a modelist engineer now, building on the Plaid AI team. And this talk is going to be less about best practices and perhaps more around different things to consider as you try to figure out how to move from using flimsier ancient systems to something that's a bit more reliable. Something I learned when I was a PM
00:01:07 was asking the right questions and making sure that certain things were correct even as I delegated technical executions. So this is the system I built at a high level where essentially Vercel gets tons of field data, calls, phone calls, support cases, Slack threads, mentions on Twitter, that's actual people just submitting feedback from our field. And we wanted to find a way to connect that to our product feature request. So I built a system that both runs through all this data,
00:01:40 Connecting signals to feature requests
00:01:44 connects it to our existing feature requests and then also pushes it out to HTTP. So this is sort of what it looks like. I had Codex generate a fake image, so this is not really our actual data, but there is a top 10 view, you can go into the features and see the accounts that are linked. Additionally, kind of hard to see from the back, but you can see how many sort of data evidence points have been accruing week over week as different things get associated to the features.
00:02:15 And every week, 22 of our product teams get these reports that essentially provide sort of an overview of all of the different data signals. So. So, for this example, 40 calls were analyzed, eight of them are relevant, 50 support cases, six are relevant, and then deep links to the actual evidence that support these, that summarize these. So, I find that when people try to make these systems, you sort of end up with these two counts, where one, maybe you got it working with Claude Code on your laptop,
00:02:48 where I was actually talking to one of our customers, and he was trying to figure out how to use RStack to go from having his laptop, essentially, running Claude Code, where he couldn't go on vacation because his workflows would essentially stop if the laptop closed, or the other side, where maybe you're using,
00:03:05 Retrieval costs and design
00:03:06 like, this is, I probably shouldn't show the example. Let's say you're using, like, Notion Agents, and you're, or Crockbot, and you don't have sort of the engineering sense of how to lower costs. I've heard examples of people dropping, like, 40K a month on a single person's workflow, just because they're blasting Office, and it works, but it's super expensive. So, how do you make it repeatable, expectable, and affordable? So, when I approach this problem, you know, for the most part, I largely work backwards from the end point.
00:03:42 I spent some time at Amazon, so working backwards has really sort of ingrained in me, where you can use your agents, you know, codex, podcode, whatever, pull as much data as you can, and then use it to start molding, or generating an output that looks nice. Once you get something that looks good, you can sort of start working backwards, you know, interviewing your agents on, like, how did it actually get there? What data did it use? And you can start to figure out, like, how you can piece these things together.
00:04:11 So, for this, there's so many different data points that I have to factor into these pipelines, where, you know, the goal of data point to a feature request looks very different across these, where, say, in gong calls and support cases, there's a lot of context that you can actually use to create a relevant linkage, whereas, you know, a tweet or website feedback, we need better visibility, it's pretty clear, but there'll be some that's just like, G0 sucks, or B0 is super expensive, but this doesn't work,
00:04:45 and those signals can be quite sparse. So, actually inspecting the data and seeing, you know, how much signal you have is really important to figure out, like, what kind of prompts we need to take those linkages. This one is interesting. So, this problem is, can be very expensive. Like I was saying, we're literally going across all of the data points that we get to make these linkages. So, going quite broad, it can be very expensive to run LLM calls over everything. I think this changed a bit now
00:05:20 with the type-safe JEV model, but I'm really excited to get access to that internally, where I was doing things to narrow things down. So, using vector embeddings to go from,
00:05:30 Evaluation and tuning
00:05:31 let's say there are 300 total candidates using semantic retrieval to test out, let's try doing the top K of 100, looking at that, seeing that only the top 25 usually gets about 90% of the relevance. So, I'm going to have my top K be like 30 and have pretty good coverage and have dramatically lower costs than doing the full 100 candidates. Let's see. Another thing that came up as I was thinking through this problem was feedback.
00:06:08 What was I going to say? Because data signals are not objectively relevant to everybody, like the field, product managers, engineers. They didn't really agree on what signal was actually useful for prioritization. Just trying to get any ground truth, I had a bunch of people actually just label the signals against feature requests. I found that honestly, it didn't really matter too much. There was not a lot of agreement. So, for this, I ended up not really worrying too much about having the perfect eval or having the perfect classification.
00:06:43 We largely pulled back. So, it's more around, is this data signal relevant to a product team? Would it be useful to see? And on top of that, adding sort of an unlinked flow so that if there are bad linkages, we can both simply remove them and that can feed into our classification moving forward. But this can like tune more towards what people actually are showing in usage rather than sort of some arbitrary ground truth of what is relevant or not. Another thing too, Resell workflows is quite nice for this.
00:07:16 So, when you're trying to do these applied AI sort of classification pipelines, saving things in batches and keeping them safe as sort of checkpoints is quite useful. When I was starting to do this work in the early stages, I would have these big workflows and I would end up accidentally scrapping sort of intermediate outputs, but saving things is not only good for both the audibility of your work, but then also you can sort of regenerate checkpoints. So, like when I was doing the reports into Slack,
00:07:47 being able to quickly use that same context and just generate a new report to see if this output was more palatable to our PMs and engineers was very useful. Lastly too, a lot of this too is like loops with agents. So, generate the full end-to-end pipeline, generate the classification, generate the report, have another agent actually look at it, look at the data and inspect whether or not the full run is actually useful. That was very helpful to scale this across all of our product teams where I just don't really have the time
00:08:23 to inspect all of that data, but throwing Astra at this was amazing. You definitely don't need to go that far with your models, but scaling horizontally was very useful knowing. Let's see. Here's sort of a high level of different components that we use for this. So, for SELPRON to trigger the workflow every week, workflows themselves for durable execution, which is quite nice when you're running like thousands of LLM calls and you may have timeouts or failures.
00:08:58 Our EVE agent framework is what delivers the digest into Slack. So, people can actually just reply to the report and immediately sort of dig into those data points. And then, AISDK and AI Gateway, you can kind of use all, get all your AI stuff with Resolve. And I'm running pretty short on time. So, let's see. I think lastly, definitely lean into the proto curve of models. You got so much work out of 5.6 Luna
00:09:36 as opposed to larger models. And that's my talk. And by the time, if you're interested in connecting, this QR code takes you to my X account. Thank you. Thank you.
00:09:50 Questions
00:09:50 Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Yeah. Any questions? I didn't start. What was the, what was like the biggest surprise in site product feedback with two tagging two accounts? Like what was kind of unexpected and difficult where we looked at that? That there wasn't really like Jeff. I'm sorry. Sorry, just repeat the question. Oh, the question, what was the most surprising thing about it? I think it was just really that people just disagreed so much where it was less about getting accuracy
00:10:29 and more around going to the different teams and getting involved in teams, you know? Yeah. So I'm just trying to settle it. Yeah. I just don't want to start it. Any other questions? Yeah. Do you think there was negative feedback because it was so AI forward or like AI committed or not like, like you've been a person? Yeah, that's a good question. I think you can repeat the question. Yes. The question, was there like a negative reception because it was so AI forward versus like even?
00:10:58 Yeah. Like I didn't go, I didn't launch any of this stuff without like talking to a bunch of PMs and like getting their leadership on board. But yeah, I did not send those reports without getting behind specific people and also like showed how like Kevin gave feedback and Elizabeth did this. So it was like, I'm sort of on the team. Time for one more question. The other ones. Yeah, go ahead. So for evaluations for your AI workflows, you had one slide where you spoke about using
00:11:33 like the third one is on these two judge of the ultra-professional pipeline. So we wanted, are there any other specific or counterintuitive facts you want to? Yeah. The question is, are there any other counterintuitive things for evaluation? Not too much. I mean, when I was designing this, I would just really use like the smartest, biggest model I had and sort of just build that down and drop rails. But yeah, pretty much just up with larger models for the design process.
00:12:09 Awesome. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. Thank you. I appreciate it.
