Find more data elsewhere
Marco says: “Use multiple data sources, not just SEO data, to make decisions.”
What's the difference between the alternative data sources and the typical data sources for SEO?
“Most of the time, people use Search Console, which is our default option for several reasons. For one, it's the only first-party tool that contains Google queries, which is the dominant search engine. People also rely on crawl data from tools like Screaming Frog and Sitebulb, or they scrape websites themselves. There are other third-party tools as well, like Majestic, Semrush, Ahrefs, etc. That’s the traditional suite.
Some people have a decent knowledge of GA4, which is used for other data and tracking conversions, events, and what happens inside the website – which is not strictly SEO, of course.
However, there are data sources other than GA4 that are also super important and can help you. For example, in your CMS (Content Management System), you can get metadata about your articles and pages that can be useful. For example, the classification of a page or even the tags or categories you assign to a given page.
Also, your CRM. If you're doing B2B, you can’t exactly connect your SEO data, like queries, but it gives you the opportunity to better understand how your website contributes to your leads, because there are connections with Salesforce and HubSpot, if you're using GA4.”
What page classifications can you get from your CMS, and how do you pull that through in order to utilise it within your analysis?
“If you are a publisher, you might have a series of articles labelled ‘GA4’, because all of those articles are about GA4. That is already a tag, or it can also be a category. It depends on the structure of your website.
This information is important, and it can be pulled via APIs, or you can just scrape it (which isn’t optimal, but it can also be done that way). This is useful because you can use it to classify your performance by sections or groups, which is more expressive than having a bunch of pages without any classification.
Speaking of which, you can also use content plans. Imagine these as a content inventory, where you have a list of pages with your URLs, and then you have columns classifying and labelling them. This exercise is very useful, especially for small websites, when you're just starting out, because you are able to label and classify your pages manually, which is more reliable than machines. That way, you will understand where you are going.
For my personal content, I use a free tool called Obsidian, where I create a content plan that contains every article and some metadata. If I want to analyse this data and compare something like organic traffic in GA4, I can just combine the tables and see which groups are performing better.”
Do you classify your content by stage of the buyer journey, by category of content, or by the type of content?
“All of them, depending on the website. Classifying based on where the user is in the funnel is quite high-level, and you would need to go into some detail.
If you are talking about a commercial page, it's not transactional. If you have a product page, the intent is that people are most likely going there to buy. If it's an article, it's a bit more complex because not every article or piece of content is super high-level, and it doesn't need to be.
A case study is not comparable to an article about how to paint rocks. They are on different levels. Normally, I try to be as detailed as possible, as long as I'm sure this data will be used in the analysis.
For example, I tend to use performance indicators with a range. I always try to label pages or content with a range, based on some metrics, to understand more or less how they are performing.
You can also use categories that you make up, or instead of classifying the article by topic, you can classify it by author, and this data is your CMS. You can also classify them by stage in the buyer journey or by audience, if you have multiple audiences.”
What data are you typically looking for in a CRM?
“CRM data is not SEO. It's another topic that you have to study. If you are working consistently in B2B, it's highly recommended that you at least understand the basics of the jargon.
You can’t actually build a direct connection to say that specific queries brought a specific amount of traffic. It's a little bit more complex than that. However, it's important to understand the connection from GA4 to Salesforce and HubSpot, for example. If you can set it up (it's not always possible), it allows you to better understand how you are affecting the sales funnel.
Even if you only consider Salesforce data, without connecting it to any other web data, it's still important to understand because, if you're doing SEO or any marketing activity, you need to know how the sales funnel works. If you don't understand that, you might think that people will buy your product in two days. However, if you are selling data centres, you don't sell a data centre project in two days. It takes one or two years.
Once you are familiar with your sales cycle and the mechanism that you have in Salesforce, HubSpot, or any other CRM, it’s much easier to do SEO – or marketing in general. You will understand that some content or some material will not perform properly, or is not beneficial.”
Is GA4 where you prefer to analyse your data, or do you prefer to use an alternate platform like BigQuery?
“Normally, you do it inside a data warehouse because you need to store your data, you need to have it in one place, and so you can monitor the cost. That way, everything is under one umbrella. You would do it in BigQuery, or any other warehouse, depending on the company you are working for. It's more ideal.
The real challenge is often the engineering part, like how to piece things together, save money, and make it performant. However, this is not an SEO consideration. Data warehouse solutions are extremely beneficial because you have everything in one place and you actually store the data, which is super important on the website.
With GSC and GA4, the APIs they offer are limited. GSC does not give you all the data, only the last 16 months. GA4 data is also limited and highly processed. You want your data to be as raw as possible. Google has some nice connectors that you can put directly into BigQuery, which is very convenient.”
How do you determine which data sources to prioritise and focus on initially?
“I have an approach that is very simple: it's based on the questions. If you have some users, some goals, or some critical business questions that you want answered, you see which data source can help you answer them. Those are your critical data sources.
A mistake I often see is when people use GA4 to report on financial data. The issue is that GA4 is limited for many reasons (for example, consent mode). Also, GA4 is not supposed to be an accounting system. That’s not its goal. In those cases, I would recommend using the system where you are recording your transactions. With transactional data, you need to be 100% sure that it is correct, for both accounting and legal reasons. In that case, the accounting systems where you are keeping your transactions are your critical data sources.
Your SEO data or web data is also important, but depending on the type of business, I wouldn't say that GSC is more important for a B2B traditional company selling tangible products than your accounting system.
GA4 and GSC are super important if you are an aggregator website, and most of your revenue comes from the digital world. Also, if you're in e-commerce, of course. Then, these are quite important – especially GA4, because GA4 tells you what happens inside your website, if you set it up properly.
This doesn't mean there aren’t other sources that are less critical but nice to have. For example, the plans or inventories I mentioned before, where you can add metadata.
Depending on the type of business, there are variations. If you're a publisher, you are not selling anything. You don't have any leads. If you are in B2B, you have leads, but you don't necessarily have transactions like you would in e-commerce. They need to go through sales first. It's a different process.
Depending on the type of business you have and the critical business questions you are answering, you can identify critical sources. Usually, these are the ones where money or opportunities are recorded. Web data can also be super critical in some businesses, but SEO data is not what’s most critical to the company.”
How much time should you spend on understanding what to do from a data perspective, and understanding platforms like GA4 and BigQuery?
“The answer here is heavily reliant on your role. If you are a lead (like a Head of SEO) and you have more political power, this is easier to achieve. If you are a specialist or you don't have much say (which is completely normal), you might not have as much control because there are silos, there is gatekeeping, and people don't want you involved in some things.
However, if you are on the lower end of the scale, you can still understand the business model, even if you don't know all the data, as a way to get your foot in the door and ask for access or understand a little bit more. If you are more of a technical professional or you’re on a higher managerial level, this is much simpler. You can actually start connecting the dots and cover topics that are a little bit broader.
For example, instead of SEO, you could consider acquisition or growth. In my opinion, this topic is easier to connect to other areas of a company, and there is less stigma attached. If you want to understand your impact or the value you are producing, you at least need to have an idea of how these systems work. How do you know that your work has produced results?
GA4 is not exactly reliable, so maybe you want to check the source systems. Then, attribution is quite complex, but it’s important when it comes to understanding how SEO has incrementally contributed. What is the unique value that SEO added, excluding the effect of other channels?
All of these considerations are a little bit more technical, and they require a lot of work. If you are just starting out as an SEO, and you have no clue about data, you can start with the business model. How does the company make money? Where are they keeping this data? Do they actually have data or not? Sometimes, they won't even have some of the data that you want.
After that, you can start piecing things together, troubleshooting, and thinking about how you can connect or use other data sources. For beginners, you must know Search Console, Screaming Frog, Sitebulb or any other crawling and third-party tools, because you need to use them for research or backlink analysis anyway: Majestic, Ahrefs, Semrush, etc.
If possible, study GA4, because it's extremely important to justify what happens inside the website, and to track conversions, events, etc. After you have this foundation, you can start with BigQuery, the exports for these tools in BigQuery, and the raw data. Once you're more familiar with all these topics (which are already outside of SEO), maybe you can also have a look at Google Ads and uncover other data sources, like SAP or Salesforce, depending on the company.
You can get a bit broader, but only after you have mastered your own area.”
Marco, what's the key takeaway from the tip you shared today?
“Don’t rely on SEO data alone. If you only use SEO data, you lack the business perspective.
It's always good to consider the whole overview of the business and the limits of your position, your role, and your tasks. In particular, avoid auditing websites only with Search Console. It can lead to dramatic results.”
Marco Giordano is a Data and Web Analyst at Seotistics. Find out more over at Seotistics.com.