[Alisa Scharf] {00:00:02} Hello, everybody, and welcome to the September edition of The Signal. My name's Alisa Scharf, I am your host today, my co-host, Nick Haigler, is here with us as well, and They're calling this, the webinar, so nice, they recorded it twice, because… and actually, we only recorded it once. Yesterday, we had this live edition, we went through all this great material, we had a little technical difficulty, so we're re-recording, in case anybody watched yesterday's episode. We're gonna try to cover all the same information, all the same content, but you might notice just a little… a few little differences. Alright, let's get right into it. So, here's what we're going to cover today. We've got a lot of good stuff planned in the agenda. First of all, Nick is going to talk through some research that we've conducted in partnership with Trustpilot, which answers a question we get a tremendous amount at Seer, which is. How should I think about reviews? Of course, reviews are important for customers, but how should I think about reviews in the context of search marketing and geo and all of these things? So we're going to answer that question. Uh, the second topic we're gonna cover is on prompt Frequency, which I know sounds very in the weeds. What we're talking about here is how many times do you need to, uh, trigger a prompt in order to get results back that are reliable enough to make business decisions based on. I actually read something really interesting last night from Mike King, and I didn't read the full post, but the idea… I've got it in my Readwise, ready to read, but it's a long one. But I skimmed it, and I think the idea he's pushing here is that You know, you're never going to get to perfect accuracy in this world, so don't try. If you're on this mission to get 100% accuracy in your prompt tracking, you're always going to fail. So, we're going to take an angle of that conversation today, and I encourage you to check out Mike's post as well. Before we get into these things, we have a very fun new segment in the Signal, and I've got a little soapbox I want to get up on and share what I've been thinking about, so let me just pull the slides for a moment. What I want to address is this… this idea, the national coverage we have been getting on the AI boom, and the negative sentiment around the AI boom. And there's been a lot in the news lately about this topic. And if you're like me, maybe you've been vocal with your friends and family about the work you're doing with AI, how it's helpful, how it's, you know, this fascinating subject. And I've gotten some texts this week from friends and family being like, I'm hearing AI is gonna cause the end of humanity, should I be worried about this? My thought is, you know, at this moment in time, no, but there's certainly things we need a bigger public discourse on, and so I'm glad that we're talking about it, but I want us to talk about this productively, and so I have some thoughts I wanted to share to that end. My first recommendation, because when things are… when things feel tough, when you kind of feel stuck in the mud, I think it can be really helpful to be like, okay, what's an action I can take? If you did not listen to the AI show this week, and that would be, uh, today, it is September 18th, so a few days ago, the AI Show had a fantastic podcast that covered the moment we are in right now. As usual, Paul Raitzer and Mike Caput from SmarterX. Uh, these are the guys that put on the Macon conference. They had a very pragmatic, thoughtful approach to how they told this story. that I think is really helpful, first of all, for everybody in our space to listen to, but it's a really helpful podcast to send to friends and family if you're getting those questions, if you're getting those thoughts, um, and you want to give people a better, more… grounded-in-reality view of what is happening, because there's certainly things to be concerned of, right? I don't want to say that's not the case, but I think we can move forward in a much more productive way than the tenor feels, um, you know, this week and last week. The other thing I wanted to highlight is one of the… one of the concerns with AI that I keep hearing is the impact on entry-level jobs. And I personally am really worried about that. I think that is a serious concern, and I think a lot of people are worried about that, and this national coverage and the fact that this is a big story is causing more and more people to kind of… Gravitate towards that problem, and yeah, this is a big deal, what are we gonna do about this? Well, AI, I firmly believe, augments experts at what they do in order for them to do more, in order for them to do better work. you still need that expertise, and this is the… this is the fear with entry-level work. If you don't have entry-level employees, how do you build those skill sets, you know, that Nick and I went through early in our careers? Doing all this work manually. enables us to do this work using AI and spot when things are wrong, spot when things are right, understand, you know, how to use these systems. So here's my recommendation. Be a part of the solution. So, there are so many opportunities for us to spend time with our community, with young adults, with young folks who need help, need mentorship. I personally know I would not be nearly as far in my AI journey if I didn't have so many fantastic mentors at Seer, and I'm constantly learning from people like Will, and people like Jordan Strauss, and people like Nick. And so I think there is a big opportunity for us to extend that… that grace to folks who don't otherwise have these opportunities. And I was really inspired to share this, because a couple of days ago, I attended this event with HopeWorks. For those who don't know, HopeWorks is a local organization, and their mission is to help young people Build the skills, build the network to, uh, to enter a field, enter, uh, an employment in tech. So, there are lots of opportunities locally, uh, within all of our communities, all of our organizations. If you're local to Philly or New Jersey, both HopeWorks and Launchpad are really, really good resources for this. So, if you're worried, if you're feeling down, if you're feeling stuck in the mud, my recommendation, listen to the podcast that I recommended. Paul does, again, a really good job of just kind of laying things out in a very neutral way, in a very not-emotionally charged way. And then number two, be a part of the solution, right? Everybody on this who's watching this has a skill set they have developed over time, so think about what you are doing to help pass that skill set on to the next generation. Alright, off my soapbox, I'm gonna hand it over to Nick to get back to our regularly scheduled programming and talk a little bit about reviews. And you are on… on mute, Nick. [Nick Haigler] {00:07:41} Thanks, Alisa. Um, so to transition us to how reviews are shaping AI responses, just for some context on the research we're about to share, Trustpilot reached out to us to help them quantify the value of their review platform in AI Search. And that led us down the path of, overarchingly, learning how reviews are shaping these AI responses. So, for some… context of the data that we're looking at, we analyzed over 800,000 responses across 4 platforms. We were looking at ChatGPT, Gemini, AI Mode, and Perplexity. And what we found is that review and trust sites are the number 2 citation source in AI responses in about a quarter of all the citations. So, one of the things that stood out to me, but I don't think it's necessarily surprising, is that review sites grow from just 3% of citations at the awareness stage, to then getting closer to that 23.5% as the users move down the funnel into making a decision. Again, I don't think that that data is necessarily surprising, but the quantification of their influence growing nearly 10 to 20x by the time users are trying to make that decision, I think is something that really stands out, and maybe something that we haven't looked into as much in the past with, um, the citations of this stage. So… At the awareness stage, these review sites are making up a small slice. They're not really being used for brand discovery, but that's where we're seeing a lot more of Reddit, at least historically in ChatGPT, and we're also seeing a lot of those editorial sites that we classified here, um. at just 3.8%. That's where the editorial sites have a much larger percent of the pie whenever we're looking for discovery. But again, as users are moving down the funnel. They're showing up, uh, the trust sites are showing up where it matters most to users. And so, just for some transparency, Trustpilot… brands that had a Trustpilot profile did account for the majority of the prompts that we selected for this research. But also, it was very rigorous and tough to find some of those brands and sites for comparison that didn't already have a Trustpilot profile. And so, whenever we were analyzing this, we found that Trustpilot did lead review platform citations across all of those four platforms that we tested. ChatGPT through Google and Gemini. So… But one of the other things that we found was we actually wanted to peel this layer back, not just look at citations, but also dig into the fan-out queries a little bit. And in this case, we found that, um… The fan-out queries and those branded mentions that we've been talking a lot about. We're commonly mentioning the Better Business Bureau, Google Reviews, and Yelp have those branded name searches in fan-out queries. Trustpilot was fourth in the amount that their brand was actually mentioned, but still, Trustpilot was gaining these citations more often in ChatGPT, even though it wasn't being mentioned as much within the fan-out queries. So I think that says something with the accessibility of the Trustpilot site, that ChatGP is able to find the information that it's looking for within the page structure and within those pages that it's searching. The second side of that is the actual site operator. queries within FanOuts. That's been another very hot topic, as ChatGP has become very much more selective on the pages and the domains that it's searching for, for any given prompt. And what we found was that over… about 23% of all of the fan-out queries use the site operator overall. The interesting piece with that is that the majority of those, about 98% of any of them using the site operator, were targeting very specific pages, um, especially on Trustpilot. They weren't searching for the entire Trustpilot domain, they were actually searching for the brand-specific review pages, the product and location pages on Trustpilot, so they were getting very much targeted. They knew URL… they knew Trustpilot's URL structure in this case, so that I believe that had Some of the, um… Was some of the reasoning why Trustpilot was getting cited as much as they have been. The next is that trust… Uh, trust signals are the most featured, or these are the trust signals most featured across responses. And whenever review content is pulled in, star ratings are the most used element, and about 60% of those responses that cite any trust or review profile. The rest of these elements are pulled through are still significant, because I believe that this just highlights the value of social proof. that would otherwise just look generic in AI responses as users are reading them. And I think there are two of these primary types of users that are reading some of these responses. I think the first is those users that are just skimming through the AI responses. We all know that, unless you have some custom settings, that sometimes these AI platforms and AI responses can get pretty lengthy, and so I believe that this and the visual of just some of the star ratings, the visuals that come through are just going to help your brand, or help the brands that you're working with, stand out a little bit for those skimmers. But then, of course, the ones that are probably even more focused on making a decision, and some of those that have higher purchase intent, for those being thorough with their decision, you have to have these trust signals, other… that will make a user more confident in the selection that they are making. Anything from stating the star rating to, uh, providing that quantified proof. Saying that, well, this is from over 1,200 reviews, and then, um, highlighting some of those, um, trusted adjectives, as well as just, um, some of that verified reviews and trust language that we saw. Each of these platforms pulls and highlights reviews in a different way. We found ChatGPT had the highest usage of star ratings, using star ratings about 80% of the time whenever Trustpilot or another review platform was being cited. We found that Gemini was the most balanced. It was using star ratings, but it was also paraphrasing the reviews that it would pull through more often than some of the others. And then Perplexity was mentioning the review verbatim much more often than any of the others. So there are… is a level of… diff… there's a different level of how these AI platforms are using their review platforms, so whenever they are citing them, that would be, um, seen and I guess, measured by the users. Another piece that I think is interesting is that review content amplifies sentiment in AI narratives. Now, the top two boxes that I have on the screen, I don't think are surprising whatsoever. we should… all probably assume that if these review platforms are getting cited, that they're going to highlight the good things and the bad things that are being mentioned in their reviews. We classified this as positive amplification and negative amplification. And this is exactly what we found. Whenever reviews are being cited, it's going to mention… if there's more, um, positive things to highlight, that's what it's going to pull through. If there's maybe more negative things to highlight, it's gonna pull through those. So it is very much… Fair, so to speak. the interesting pieces here, I believe, are the Comparative amplification and the absence Signal. The Comparative amplification is whenever it's pulling multiple different review platforms, and it's comparing what it's seeing on Trustpilot compared to what it's seeing on G2, or the Better Business Bureau, or any of these other review platforms, to provide much more information and confidence to a user So they have an overarching view of what a product or brand is about. The most interesting piece, I believe, though, is the absence signal. This is where I think could help or hurt the brands the most. If it's actively saying there's a distinct lack of extensive verified third-party reviews on these mainstream consumer sites, like Trustpilot or Better Business Bureau. That's something that whenever a user is actively making a decision and they see that there's a lack of trust, there's a lack of authoritativeness, that it's very much going to steer a user in the opposite direction if they're not willing and want to learn more about that brand on their own. They're actively being steered away from a brand. And… Within some of these other brands, we had a set of comparison Prompts that were set up. Let's compare brand A versus brand B. And what we found is that the brand with more reviews gets its evidence pulled into AI responses almost two times as often as a brand with less reviews. Now, this could be a brand comparison where one brand has 1,200 reviews, and the next brand has 1,000. Still, very thorough and very, um, in-depth review site, or review, um… Inventory for a brand, but just because it has less than the brand that it's getting compared against, its evidence isn't going to get cited as much, or its evidence isn't going to be shown to the user as much to provide them that social proof and that confidence. And here's an example of what that actually looks like. We have a prompt set up for which is better, brand A or Brand B in Miami. Now, these were two separate recruitment offices, and what I want to highlight is the difference that these brands are being represented within the same response. Brand A, appears to have a stronger public reputation, large number of recent reviews, around a 4.9 overall on Trustpilot, consistently positive feedback for the Miami office. It's… this brand had, obviously, the most reviews in its comparison, and then within that same response, right below where brand A is getting complimented and being shown to the users as, hey. These people know what they're doing. They're, um… everybody else seems to really prefer them. Right next to that, you see Brand B, and Brand B has much less publicly available feedback. So it's not even really quantifying any of this, it's just saying, comparatively, comparing… compared to Brand A, this one has much less. Which, so you're actively getting compared against and creating some of that bias in these AI responses. It's also highlighting that there's really not any independent resources available to understand this brand at a deeper level. All of this information is coming from the company's own website. So, it's really not highlighting any of this proof that users are looking for. So, as some next steps, I think that all people working with, um, in-house or working with other brands should test some of the comparison content prompts to see how your brand is being presented, and where that information is stemming from. We did this with Seer, and we found that one single review that was from over 5 years ago was driving a large percentage of our brand perception in ChatGP ChatGPT, and it's something that we wouldn't have known unless we actively started doing some digging and understand how we were being framed. For smaller brands, I think piggybacking on the authority of Trustpilot and other Review sites that are specific to your space is the single most effective way to increase your trust signals, where ChatGPT, some of these other platforms may not directly go to your site. You may not be actively being included in those site operator searches within the fanouts yet. So, be visible where Where they are searching, and piggyback off of that authority. for larger brands, I would keep the brand A versus brand B comparison in mind, even if you have a very thorough review process set up. If you're being compared to someone else in your space, and you know that that's who you're commonly being compared against, check out their review sites, see how they're being listed, because that's often going to be representative or compared against how you are being represented. And then the last thing I want to mention is that I believe there's another level of reinforcement. A lot of times we think about signal reinforcement on trying to get onto other websites. We have all of this great information on our own website, and we want to push that information onto the third-party sites, like Reddit, like earned media and PR sites, so that that gets highlighted more to users and to AI. I think this is a little bit of reverse, where you can see what signals are already existing on Trustpilot, on other review sites. See where your brand is getting bragged about. If it's on Trustpilot, maybe there's a certain space on your own site that you should you know, reinforce that content. Um, Seer, we created our testimonials page, um, just for this sole focus, and we're already seeing that it's being cited, um, and referenced in AI responses, even though it's a very new page as well. So… I think this can be completed through having that sole dedicated space on your site, or using some of the widgets that these review sites offer as well, to highlight those signals. [Alisa Scharf] {00:21:14} Awesome. Great stuff. Thank you, Nick. And we're gonna talk more about reviews, so stay tuned, we've got a bunch of Q&A questions. We'll hit on reviews. Uh, but, alright, I'm gonna jump in. We're gonna switch gears to talk about Prompt Frequency. So again, we're talking about AI search here, and platforms like ChatGPT, or Claude, or Gemini. And tracking and measurement looks different for this space than it does for SEO. Remember, we have decades of experience and the maturity of tools and technology supporting our ability to measure SEO, and even that's not a precise measurement, a precise science, but we've got business trust, right? We've got executive trust in that data, and we're not there yet with AI search. And this is one example of why I believe we're not yet there. So, Nick, you can go to the next slide. We ran a little test with just under 100 prompts. And the question was simple. If I run this prompt one time, how many brands do I get back? If I run the prompt 3 times, or 5 times, or 10 times, how many more brands enter the competitive set? And what we found is that one run will show less than half the field. So, in other words, you do a prompt, you run it once, and maybe 5 or 6 brands appear. If we run it 10 times, all of a sudden we're up to 12 or 13 brands. And so if you are representing one of those brands that didn't make the cut the first time. you may be reporting on data that is not entirely… I don't want to say accurate, because as I said before, none of this is accurate, but at least directionally sound. So a lot of the tools in this space are new, right? there's some players who have been in this game for a long time, right? Like, folks like Conductor or BrightEdge, but there's a lot of new players to this space, and so I think the… the methodology still needs… still needs work, and that's collectively, as our industry. This is something to address. So, yes, exactly, go to the next slide. So, this is another cut at that same data. three runs. So, run the same prompt three times, and you're gonna get a little more than 70% of the brands ChatGPT is ever going to name. And apply this same probability to the experience your customers are seeing, and all of a sudden, this starts to feel pretty directionally sound, right? It's not comprehensive, it's not the end-all, be-all, but we know there's a lot of nuances to tracking prompts, right? The customization is far more rampant, the memory implications, what tools or data sources you have connected. So the goal is to get a number that you can go to your board, you can go to your executive team, and say, look, this isn't perfect, but if we measure it this way, we benchmark it that way, this is going to tell us if we are winning or losing. So if you go to the next slide, an interesting element that we found with this research, and we ran a sample of prompts that represented a really wide range of industries, like categories. We did that intentionally, because Nick and I both had a hunch that, I think depending on what category you're in, you're going to need a different frequency, or a different volume of Of prompts to track. And so what this tells us is if you are working in legal services, doing marketing for beauty and personal care, health and fitness products, energy and utility, there's far more volatility in that space. So if today. you're tracking prompts once, I think the message is everybody needs to rethink that. But if you're in these spaces next to these pink bars, you might need to be tracking prompts 3 times, 5 times, 10 times in each instance. this has implications, right? Because that's… that's not free. This is not free data. We don't have a data source like Google Search Console, where we can say, okay, let's use this as the free source of truth, and we can complement this with another paid tool. So I understand there's budget implications to this, but again, if we want to get a board or an executive group to trust this data, we've got to at least address some of the elephants in the room. Now, the positive story, though, is if you are in a space like restaurants, food and grocery, home improvement, travel and hospitality, and some of these other industries that you see at the bottom of the list. I still think you need to track your prompts more than once, but your frequency maybe would be fine if you… if you're tracking all of your prompts 3 times, right, instead of 5 or 7 or 10. So that's good news for those groups. Can go to the next slide, please, Nick. So, covered this in the last one, but the big takeaway here, based on this research, and there's more research Nick and I are going to be doing, um, based on this research, the floor is 3 runs per prompt. Again, I think… I don't think there is a space out there where you can reliably just track your data with one prompt. One prompt one time, and say, yes, I know what's happening in this space. And what you're trying to get to is a really reliable share of voice or share of market metric. This is the metric I think we as an industry need to be able to say, look, we are planting the flag. There is value here. Even if you can't measure the click. from the platform all the way to analytics that says that there was a conversion associated with this user, there is still value in being mentioned, because we know that the clickstream of this whole channel is going to be broken. It's going to be really, really worn down, and there's going to be a lot of missing pieces that make it look like Actually, there's only a little bit of traffic coming in from ChatGPT, so we don't have to worry about this yet. What you might also see is, oh, it's strange, our direct traffic is going up, our branded paid search traffic is coming in higher and converting more efficiently, our general web referral traffic is up. So there should also be measurable signs, right? The goal is not just get a mention, the belief is that if you're getting a mention or you're getting a recommendation. It is going to have a business impact. My POV, though, is that we're not going to be able to tie those things together like we were 5 or 10 years ago with SEO, and so we need a better approach. That's why the integrity of this measurement is so important to us. If you go to the next slide, Nick, this, I think, is a bright spot and something really interesting. I didn't expect to see as a result of this data, but if you happen to be a brand that tends to rank number one, rank number one, I know I'm, you know, taking liberties here, but if you're the first brand visible for prompts that you are tracking that's relevant to your audience, that is a really good place to be, because there's a level of stickiness there. That is not really common in this industry, right? So, often, part of the issue people have with the volatility and the probabilistic nature of measurement with this channel is that you could ask the same prompt 10 times and maybe get the same 3, 4, 5 results, 3, 4, 5 brands. They're all gonna be in different orders, though. What this data showed is that's not necessarily true for the brand that is mentioned first. There's a stickiness there where 75% of runs repeated the same lead recommendation. I will say, I think if you are one of these brands, it is very rare that this is a result of something you have intentionally done for a geo program, there's a bit of, hey, we were born on third, and now it's really easy for us to score. That's fine, because if you were born on third base in this world, it's because you were doing things that legitimized your place as the market leader, right? So take credit for it, certainly, but I would be cautious of saying, oh, because we did these things for us to build citations in 2025, we've got this visibility. I think that that may not be the case. Okay, let's go to the next one. This is interesting, too, because as part of this research, we were tracking prompts that were meant to reveal specific brands, and there were some spaces where no brands were revealed. So if you are working in spaces with hyperlocal services or medically sensitive products. just be aware that this is a space… I would think of it kind of like Google's Your Money, Your Life concept, where there's just gonna be a different set of criteria for what the model is willing to recommend. And so, like anything in this space, it's gonna be really hard to trust those general best practices. You're gonna have to pull your own data and see what happens for the prompts that are most relevant for your brand. All of this maps back to our core ethos here, and what we really want to imprint upon our audience is this idea that this channel is very impactful for certain audiences, many audiences, in fact. It's not always going to be measurable through your analytics platform. I think there is still tremendous value in being seen, and being mentioned, certainly in being recommended, and there is a difference between being mentioned and recommended. There's also tremendous value in being believed. This is a really interesting element, too. A lot of people who are working in GEO today came up through SEO. In SEO, you don't really have to worry a tremendous amount about branded search for the vast majority of organizations, because if you want to rank for something branded. create a piece of content on your website targeted to that branded term, and Google's gonna say, yeah, you deserve to rank for it. This is a very different world, though, where it's more about what others are saying about you. Case in point, Nick's research in the first part of this presentation. And so, what you may find is you have invested a ton in your brand, or in a rebrand. And those… those facts, those elements, those details that are really important for your audience to understand about your brand and your product and services, it is far from a given. that ChatGPT or Claude or Gemini really deeply understands that. So we've written a lot about brand canon and the work there, but I think that is a really exciting, kind of emerging discipline, sub-discipline within this space that deserves to be spoken about more. And then lastly, we want to be chosen, right? We want our customers to be able to find us, access us, convert, etc. You're all good, Nick. That's the fastest one. Okay. Three things to do this month. First of all, we recommend that you analyze what review platforms are getting cited for your brand's evaluation prompts. We recommend you claim your review profile and monitor those themes and sentiment. they're likely being pulled through into AI responses, even if you think you're in an industry that this doesn't… this… this doesn't matter in. I mean, like Nick said about Seer, we dealt with this. This is… we… We… typically, uh, build our pipeline and find our clients through word of mouth, through direct referrals, through our thought leadership. And so things like online reviews are like, eh, this doesn't really matter at a tremendous amount for our business. these things are getting pulled through into AI search, though, and our audience is heavily, uh, adopting AI search, and so all of a sudden, it's like, yeah, this is something we've got to care about. And then lastly, run that 10 times baseline burst. So run those 20 or 30 prompts that you think matter most to your business, run them 10 times, and classify those prompts to see where the results are pretty locked in, where the results are mixed, where there's volatility to be aware of as part of measurement. Alright.