What is a server log and how it can help your SEO efforts
Ever heard of a server log but not really sure what it’s for? Or maybe server logs sit in the domain of your developers or IT specialists. Well in reality, there are a lot of opportunities for improving your SEO performance by analysing these logs. So we’re going to look at server logs and how they can help you get a deeper understanding on how Google crawls your site.
What is a server log?
Server logs are a collection of text files automatically created by a server documenting activity. These are typically used by developers to help monitor the health of your website, debug any errors and watch out for security issues. From an SEO perspective, we are normally interested in a specific set of logs, known as ‘access logs’. Access logs record every request made to your website, information about the requester, and the response.
157.25.182.201 – – [24/Jan/2022:13:55:36 -0700] “GET /about_us.html HTTP/1.0” 200 2326 “-” “Googlebot/2.1 (+http://www.google.com/bot.html)” “www.example.com”
For example, in the log entry above, we can see:
- 157.25.182.201 – the IP address of the requester.
- [24/Jan/2022:13:55:36 -0700] – the date, time and time zone that the request was made.
- “GET /about_us.html HTTP/1.0” – the HTTP method, GET , the resource (page) that was requested, /about_us.html , and the protocol used, HTTP/1.0.
- 200 – the HTTP response state code.
- 2326 – the size of the server response in bytes.
- Googlebot/2.1 (+http://www.google.com/bot.html) – the user agent of the requester.
- www.example.com – the referring URL.
How does a server log help with SEO?
From an SEO perspective, server logs provide us an unparalleled insight into the exact traffic coming to our website. We can use these insights to understand how a search engine perceives our site.
For example, understanding crawl prioritization, is a domain where log analysis shines. Ideally Google’s bot will try to crawl all pages available and re-crawl them based on known URL patterns. However, if your site is on the larger size, Google may not have the crawl budget to revisit every page. Google’s bot will make decisions on which pages to re-crawl based on its crawl prioritization algorithm.
When a site isn’t being crawled frequently enough, updates to content may not be picked up quickly. This will limit your organic search rankings. By analyzing your logs, we are able to identify which of your pages are being crawled, how Google is prioritizing them, and compare that with how your pages are ranking. We can also tell how long it takes Google to crawl the entirety of our site. This could help to inform decisions about which pages to exclude from Google’s crawling.
Another area of analysis is the web servers HTTP response codes and response times. Reliable, highly available, and responsive websites are beneficial for search engine rankings. Our logs provide the necessary insight for every request versus Google Search Console’s limited sample. We can check for 200 response codes and identify any major service side errors through the volume of 500 response codes. This is especially useful when we set redirects, (300 response codes), as we can filter our logs by page requested to check that our redirects are working as intended.
An additional advantage of using server logs is that we are able to filter out irrelevant entries, such as from fake search engine bots, from the legitimate requests we care about. Though there are a range of security products that can block bots, some traffic may still get through. This traffic can fool us into thinking that our site is being crawled more frequently by legitimate bots than it really is. Filtering out this noise from our data allows us to get to the heart of what’s really happening. You can also filter log entries by search engines: useful if you just wanted to see requests from Bing! or Baidu.
How do I analyze server log files?
The first step in analyzing our logs is to get access. We need to ensure that our webmaster has enabled logging and that the logs are being stored and kept safe for later use. Many businesses are using log aggregators, such as Prometheus, to stream their logs from multiple web applications simultaneously. It’s important to check with local data regulations, such as GDPR, to make sure that the data you’re retaining is inline with the policy.
Once you have access to your logs there are many tools available for ingesting and analyzing them. Screaming Frog SEO Log File Analyser is an example of a log analysis tool built specifically for SEO. Tools such as this provide an easy way to view crawl frequency, identify crawled pages and find broken pages. Google also provides their own tooling within the Google Search Console. Here you will be able to find crawl response status codes, crawl purpose and Googlebot type. It’s worth noting that this data only represents a recent snapshot, the last 90 days, and isn’t as comprehensive as the data available to you in your logs directly.
The real benefits of log analysis come over time. By continuously logging you can hopefully see gradual improvements. This is especially important if you’re undergoing a site migration or changing a site’s structure. By retaining the logs and comparing against recent data you’ll be able to review SEO performance over a substantial period of time.
Overall, performing log analysis allows you to really get under the hood of what’s happening on your website. As the old saying goes, knowledge is power, and web server logs are often an untapped and unappreciated source of knowledge for the SEO professional.