I haven’t filled you guys in on the plumbing case study in a bit. There hasn’t been much change once I had gotten into the teens. Most of the keywords have settled around the 14th to 24th spots, with a … Continued
The bird has landed, and by bird, I mean the MozCon 2015 Video Bundle! That's right, 27 sessions and over 15 hours of knowledge from our top notch speakers right at your fingertips. Watch presentations about SEO, personalization, content strategy, local SEO, Facebook graph search, and more to level up your online marketing expertise.
If these videos were already on your wish list, skip ahead:
Before you start reading, I want to say that I am not an analytics expert per se, but a strategic SEO and digital marketing consultant.On the other hand, in my daily work of auditing and designing holistic digital marketing strategies, I deal a lot with Analytics in order to understand my clients' gaps and opportunities.
For that reason, what you are going to read isn't an "ultimate guide," but instead my personal and practical guide to content and its metrics,filled with links to useful resources that helped me solving the big contents' metric mystery. Ihappily expect to see your ideas in the comments.
The difference between content and formats
One of the hardest things to measure is content effectiveness, mostly because there exists great confusion about its changing nature and purpose. One common problem is thinking of "content" and "formats" as synonyms, which leads to frustration and, with the wrong scaling processes present, may also lead to Google disasters.
What is the difference between content and formats?
Content is any message a brand/person delivers to an audience;
Formats are the specific ways a brand/person can deliver that message (e.g. data visualizations, written content, images/photos, video, etc.).
Just to be clear: We engage and eventually share the ideas and emotions that content represents, not its formats. Formats are just the clothing we choose for our content, and keeping the fashion metaphor, some ways of dressing are better than others for making a message more explicit.
Strategy, as in everything in marketing, also plays a very important role when it comes to content.
It is during the strategic phase that we attempt to understand (both thanks to our own site analysis and competitive analysis of others' sites) if our content is responding to our audience's interests and needs, and also to understand what metrics we must choose in order to assess its success or failure.
Paraphrasing an old Pirelli commercial tagline: Content without strategy is nothing.
Strategy: Starting with why/how/what
When we are building a content strategy, we should ask ourselves (and our clients and CMOs) these classic questions:
Why does the brand exist?
How does the brand solidify its "why?"
What specific tactics will the brand use for successfully developing the "how?"
Only when we have those answers can we understand the goals of our content, what metrics to consider, and how to calculate them.
Let use an example every Mozzer can understand.
Why does Moz exist?
The answer is in its tagline:
Inbound marketing is complicated. Moz’s software makes it easy.
How does Moz solidify its "why?"
Moz produces a series of tools, which help marketers in auditing, monitoring and taking insightful decisions about their web marketing projects.
Moreover, Moz creates and publishes content, which aims to educate marketers to do their jobs better.
If you notice, we can already pick out a couple of generic goals here:
Leads > subscriptions;
Awareness (that may ultimately drive leads).
What specific tactics does Moz use for successfully achieving its main goals?
Considering the nature of the two main goals we clarified above, we can find content tactics covering all the areas of the so-called content matrix.
Some classic content matrix models are the ones developed by Distilled (in the image above) and Smart Insights and First 10, but it is a good idea to develop your own based on the insights you may have about your specific industry niche.
The things Moz does are many, so I am presenting an incomplete list here.
In the "Purchase" side and with conversion and persuasion as end goals:
Home page and "Products" section of Moz.com (we can define them as "organic landing pages");
Content about tools
Free tools;
Pro tools (which are substantially free for a 30-day trial period).
CPC landing pages;
Price page with testimonials;
"About" section;
Events sponsorship.
In the "Awareness" side and with educational and entertainment (or pure engagement) purposes:
The blogs (both the main blog and UGC);
The "Learn and Connect" section, which includes the Q&A;
Guides;
Games (The SEO Expert Quiz can surely be considered a game);
Webinars;
Social media publishing;
Email marketing
Live events (MozCon and LocalUp, but also the events where Moz Staff is present with one or more speakers).
Once we have the content inventory of our web site, we can relatively easily identify the specific goals for the different pieces of content, and of the single type of content we own and will create.
I will usually not consider content like tools, sponsorship, or live events, because even though content surely plays a role in their goals' achievement, there are also other factors like user satisfaction and serendipity involved which are not directly related to content itself or cannot be easily measured.
Measuring landing/conversion pages' content
This may be the easier kind of content to measure, because it is deeply related to the more general measures of leads and conversions, and it is also strongly related to everything CRO.
We can measure the effectiveness of our landing/conversion pages' content easily with Google Analytics, especially if we remember to implement content grouping (here's the official Google guide) and follow the suggestions Jeff Sauer offered in this post on Moz.
We can find another great resource and practical suggestions in this older (but still valid) post by Justin Cutroni: How to Use Google Analytics Content Grouping: 4 Business Examples. The example Justin offers about Patagonia.com is particularly interesting, because it is explicitly about product pages.
On the other hand, we should always remember that the default conversion rate metric should not be taken as the only metric to incorporate into decision-making; the same is true when it comes to content performance and optimization. In fact, as Dan Barker said once, the better we segment our analysis the better we can understand the performance of our money pages, give a better meaning to the conversion rate value and, therefore, correct and improve our sales and leads.
Good examples of segmentation are:
Conversions per returning visitor vs new visitor;
Conversions per type of visitor based on demographic data;
Conversions per channel/device.
These segmented metrics are fundamental for developing A/B tests with our content.
Here are some examples of A/B tests for landing/conversion pages' content:
Title tags and meta description A/B tests (yes, title tags and meta descriptions are content too, and they have a fundamental role in CTR and "first impressions");
Prominent presence of testimonials vs. a more discreet one;
Tone of voice used in the product description (copywriting experiment);
Product slideshow vs. video.
Here are a few additional sources about CRO and content, surely better than me for inspiring you in this specific field:
Here is where things start getting a little more complicated.
Blog posts, guides, white papers, and similar content usually do not have a conversion/lead nature, at leastnot directly. Usually their goals are more intangible ones, such as creating awareness, likability, trust, and authority.
In other cases, then, this kind of content also serves the objective of creating and maintaining an active community, as it does in the case of Moz. I tend to consider this a subset, though, because in many niches creating a community is not a top priority. Or, even if it is, it does not offer a reliable flux of "signals" so as to appropriately measure the effectiveness of our content because of pure lack of statistical evidence.
A good starting place is measuring the so-called consumption metrics.
Again, the ideal is to implement content grouping in Google Analytics (see the video above), because that way we can segment every different kind of editorial content.
For instance, if we have a blog, not only we can create a group for it, but we can also create
As many groups as there are categories and tags on our blog;
groups by average length of the posts;
groups per the kind of prominent formats used (video posts like Moz's Whiteboard Fridays, infographics, long-form, etc.).
This are just three examples; think about your own measuring needs and the nature of your content, and you will come out with other ideas for content groupings.
The following are basic metrics that you'll need to consider when measuring your editorial content:
Pageviews / Unique Pageviews
Pages / Session
Time on Page
The ideal is to analyze these metrics at least with these secondary levels:
Medium / Sources, so you can understand what channel contributed the most to your content visibility. Remember, though, that dark search/social is a reality that can screw up your metrics (check out Marshall Simmonds' deck from MozCon 2015);
User Type, so to see what percent of the Pageviews is due to returning visitors (a good indicator of the level of trust and authority our content has) and new ones (which indicates the ability our content has to attract new potentially long-lasting readers);
Mobile, which is useful in understanding the environments in which our users mostly interact with our content, and how we have to optimize its experience depending on the device used, hence helping making our content more memorable.
You surely can have fun also analyzing your content's performance by segmenting them per demographic indicators. For instance, it may be interesting to see what affinity categories of your readers there are, depending on the categorization used in your blog and that you have replicated in your content grouping. This, in fact, can help us in better understanding the personas composing our audience, and so refining the targeting of our content.
As you can see, I did not mention bounce rate as a metric to consider, and there is a reason for that: Bounce rate is tricky, and its misinterpretation can lead to bad decisions.
Instead of bounce rate, when it comes to editorial content (and blog posts in particular), I prefer to consider scroll completion, a metric we can retrieve using Tag Manager (see this post by Optimize Smart).
Finally, especially if you also grouped content for outstanding format used (video, embedded SlideShare, etc.), you will need to retrieve users' interactions through Tag Manager. However, if you really want to dig into the analysis of how that content is consumed by users, you will need to export your Analytics data and then combine it with data from external sources, like YouTube Analytics, SlideShare Analytics, etc.
The more we share, the more we have. This is also true in Marketing.
Consumption metrics, though, are not enough in order to understand the performance of your content, especially if you strongly rely on a community and one of the content objectives is creating and growing a community around your brand.
I usually add comments into these Metrics, because of the social nature comments have. Again, thanks to Tag Manager, you can easily tag when someone clicks on the "add comment" button.
A final metric we should always consider is the page value. As Google itself explains in that Help Page:
Page value is a measure of influence. It’s a single number that can help you better understand which pages on your site drive conversions and revenue. Pages with a high Page Value are more influential than pages with a low Page Value [Page Value is also shown for groups of content].
The combined analysis of consumption and social metrics can offer us a very granular understanding of how our content is performing, therefore how to optimize our strategy and/or how to start conducting A/B tests.
On the other hand, such a granular vision is not the ideal for reporting, especially if we have to report to a board of directors and not to our in-house or in-agency counterpart.
In that case being able to resume all these metrics (or the most relevant ones) in just one metric is very useful.
How to do it? My suggestion is to follow (and adapt to your own needs) the methodology used by the Moz editorial team and described in this post by Trevor Klein.
What about the ROI of editorial content? Don't give up; I'll talk about it below.
Measuring the ROI of content marketing and content-based link building campaigns
Theoretically measuring the ROI of something is relatively easy:
(Return - Investment) / Investment = ROI.
However the difficulty is not in that formula itself, but in the values used in that formula.
How to calculate the investment value?
Usually we have a given budget assigned for our content marketing and/or content-based campaigns. If that is the case, perfect! We have a figure to use for the investment value.
A complete different situation is when we must present a budget proposal and/or assign part of the budget to each campaign in a balanced and considered way.
In this post by Caroline Gilbert for Siege Media you can find great suggestions about how to calculate a content marketing budget, but I would like to present mine, too, which is based on competitive analysis.
Here's what I do:
Identify the distinct competitors which created content related to what we will target with our campaign. I rely on both SERP analysis (i.e.: using the Keyword Difficulty Tool by Moz) and information we can retrieve with a "keyword search" on Buzzsumo.
Social shares per kind of social network (these are available from BuzzSumo). Remember that some of these social shares can be tallied by sponsored content (check this Social Media Explorer post about how to do Facebook competitive analysis).
Estimated traffic to the content's URL (data retrieved via SimilarWeb).
Assign a monetary value to the metrics retrieved.
Calculate the competitors' potential investment value.
Calculate the median investment value of all the competitors.
Consider the delta between what the client/company invested in content marketing (or link building, if it is moving from classic old link building to modern link earning) before, as well as the median investment value of the competitors.
Calculate and propose the content marketing / content-based campaign's value in a range which goes from "minimum viable budget" to "ideal."
Reality teaches us that the proposed investment is not the same than the real investment, but at least we then have some data for proposing it and not just a gut feeling. However, we must be prepared to work with budgets that are more on the "minimum viable" side than on the ideal one.
How to calculate revenue?
You can find a good number of ROI calculators, but I particularly like the Fractl one, because it is very easy to understand and use.
Their general philosophy is to calculate ROI in terms of how much traffic, links, and social shares the content itself has generated organically, hence how much it helped saving in paid promotion.
If you look at it, it reminds the methodology I described above (points 1 to 7).
However, when it comes to social shares, you should avoid the classic mistake of considering only the social shares directly generated by the page your content has been published.
For instance, let's take the Idioms of the World campaigns Verve Search did for HotelClub.com and which won the European Search Awards.
If we we look only at its own social share metrics, we will have just a partial picture:
Instead, if we see what are the social shares metrics of the pages that linked and talked about it, we will have the complete picture.
As you can imagine, you can calculate the ROI of your editorial content using the same methodology.
Obviously the Fractl ROI calculator is far from being perfect, as it does not consider the offline repercussion a content campaign may have (the Idioms of the World campaign was organically published in a outstanding placement on The Guardian's paper version, for instance), but it is a solid base for crafting your own ROI calculation.
Conclusions
So, we have arrived at the end of this personal guide about content and its metrics.
Remember these important things:
Don't be data driven, be data informed;
Think strategically, act tactically;
Content's metrics vary depending on the goals of content itself.
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think of it as your exclusive digest of stuff you don't have time to hunt down but want to read!
In spite of all the advice, the strategic discussions and the conference talks, we Internet marketers are still algorithmic thinkers. That’s obvious when you think of SEO.
Even when we talk about content, we’re algorithmic thinkers. Ask yourself: How many times has a client asked you, “How much content do we need?” How often do you still hear “How unique does this page need to be?”
That’s 100% algorithmic thinking: Produce a certain amount of content, move up a certain number of spaces.
But you and I know it’s complete bullshit.
I’m not suggesting you ignore the algorithm. You should definitely chase it. Understanding a little bit about what goes on in Google’s pointy little head helps. But it’s not enough.
A tale of SEO woe that makes you go "whoa"
I have this friend.
He ranked #10 for "flibbergibbet." He wanted to rank #1.
He compared his site to the #1 site and realized the #1 site had five hundred blog posts.
“That site has five hundred blog posts,” he said, “I must have more.”
So he hired a few writers and cranked out five thousand blogs posts that melted Microsoft Word’s grammar check. He didn’t move up in the rankings. I’m shocked.
“That guy’s spamming,” he decided, “I’ll just report him to Google and hope for the best.”
What happened? Why didn’t adding five thousand blog posts work?
It’s pretty obvious: My, uh, friend added nothing but crap content to a site that was already outranked. Bulk is no longer a ranking tactic. Google’s very aware of that tactic. Lots of smart engineers have put time into updates like Panda to compensate.
He started like this:
And ended up like this:
Alright, yeah, I was Mr. Flood The Site With Content, way back in 2003. Don’t judge me, whippersnappers.
Reality’s never that obvious. You’re scratching and clawing to move up two spots, you’ve got an overtasked IT team pushing back on changes, and you've got a boss who needs to know the implications of every recommendation.
Why fix duplication if rel=canonical can address it? Fixing duplication will take more time and cost more money. It’s easier to paste in one line of code. You and I know it’s better to fix the duplication. But it’s a hard sell.
Why deal with 302 versus 404 response codes and home page redirection? The basic user experience remains the same. Again, we just know that a server should return one home page without any redirects and that it should send a ‘not found’ 404 response if a page is missing. If it’s going to take 3 developer hours to reconfigure the server, though, how do we justify it? There’s no flashing sign reading “Your site has a problem!”
Why change this thing and not that thing?
At the same time, our boss/client sees that the site above theirs has five hundred blog posts and thousands of links from sites selling correspondence MBAs. So they want five thousand blog posts and cheap links as quickly as possible.
Cue crazy music.
SEO lacks clarity
SEO is, in some ways, for the insane. It’s an absurd collection of technical tweaks, content thinking, link building and other little tactics that may or may not work. A novice gets exposed to one piece of crappy information after another, with an occasional bit of useful stuff mixed in. They create sites that repel search engines and piss off users. They get more awful advice. The cycle repeats. Every time it does, best practices get more muddled.
SEO lacks clarity. We can’t easily weigh the value of one change or tactic over another. But we can look at our changes and tactics in context. When we examine the potential of several changes or tactics before we flip the switch, we get a closer balance between algorithm-thinking and actual strategy.
Distance from perfect brings clarity to tactics and strategy
At some point you have to turn that knowledge into practice. You have to take action based on recommendations, your knowledge of SEO, and business considerations.
I know subfolders work better. Sorry, couldn’t resist. Let the flaming comments commence.
To get clarity, take a deep breath and ask yourself:
“All other things being equal, will this change, tactic, or strategy move my site closer to perfect than my competitors?”
Breaking it down:
"Change, tactic, or strategy"
A change takes an existing component or policy and makes it something else. Replatforming is a massive change. Adding a new page is a smaller one. Adding ALT attributes to your images is another example. Changing the way your shopping cart works is yet another.
A tactic is a specific, executable practice. In SEO, that might be fixing broken links, optimizing ALT attributes, optimizing title tags or producing a specific piece of content.
A strategy is a broader decision that’ll cause change or drive tactics. A long-term content policy is the easiest example. Shifting away from asynchronous content and moving to server-generated content is another example.
"Perfect"
No one knows exactly what Google considers "perfect," and "perfect" can't really exist, but you can bet a perfect web page/site would have all of the following:
Completely visible content that’s perfectly relevant to the audience and query
A flawless user experience
Instant load time
Zero duplicate content
Every page easily indexed and classified
No mistakes, broken links, redirects or anything else generally yucky
Zero reported problems or suggestions in each search engines’ webmaster tools, sorry, "Search Consoles"
Complete authority through immaculate, organically-generated links
These 8 categories (and any of the other bazillion that probably exist) give you a way to break down "perfect" and help you focus on what’s really going to move you forward. These different areas may involve different facets of your organization.
Your IT team can work on load time and creating an error-free front- and back-end. Link building requires the time and effort of content and outreach teams.
Tactics for relevant, visible content and current best practices in UX are going to be more involved, requiring research and real study of your audience.
What you need and what resources you have are going to impact which tactics are most realistic for you.
But there’s a basic rule: If a website would make Googlebot swoon and present zero obstacles to users, it’s close to perfect.
"All other things being equal"
Assume every competing website is optimized exactly as well as yours.
Now ask: Will this [tactic, change or strategy] move you closer to perfect?
That’s the "all other things being equal" rule. And it’s an incredibly powerful rubric for evaluating potential changes before you act. Pretend you’re in a tie with your competitors. Will this one thing be the tiebreaker? Will it put you ahead? Or will it cause you to fall behind?
"Closer to perfect than my competitors"
Perfect is great, but unattainable. What you really need is to be just a little perfect-er.
Chasing perfect can be dangerous. Perfect is the enemy of the good (I love that quote. Hated Voltaire. But I love that quote). If you wait for the opportunity/resources to reach perfection, you’ll never do anything. And the only way to reduce distance from perfect is to execute.
Instead of aiming for pure perfection, aim for more perfect than your competitors. Beat them feature-by-feature, tactic-by-tactic. Implement strategy that supports long-term superiority.
Don’t slack off. But set priorities and measure your effort. If fixing server response codes will take one hour and fixing duplication will take ten, fix the response codes first. Both move you closer to perfect. Fixing response codes may not move the needle as much, but it’s a lot easier to do. Then move on to fixing duplicates.
Do the 60% that gets you a 90% improvement. Then move on to the next thing and do it again. When you’re done, get to work on that last 40%. Repeat as necessary.
Take advantage of quick wins. That gives you more time to focus on your bigger solutions.
Sites that are "fine" are pretty far from perfect
Google has lots of tweaks, tools and workarounds to help us mitigate sub-optimal sites:
Rel=canonical lets us guide Google past duplicate content rather than fix it
HTML snapshots let us reveal content that’s delivered using asynchronous content and JavaScript frameworks
We can use rel=next and prev to guide search bots through outrageously long pagination tunnels
And we can use rel=nofollow to hide spammy links and banners
Easy, right? All of these solutions may reduce distance from perfect (the search engines don’t guarantee it). But they don’t reduce it as much as fixing the problems.
The next time you set up rel=canonical, ask yourself:
“All other things being equal, will using rel=canonical to make up for duplication move my site closer to perfect than my competitors?”
Answer: Not if they’re using rel=canonical, too. You’re both using imperfect solutions that force search engines to crawl every page of your site, duplicates included. If you want to pass them on your way to perfect, you need to fix the duplicate content.
When you use Angular.js to deliver regular content pages, ask yourself:
“All other things being equal, will using HTML snapshots instead of actual, visible content move my site closer to perfect than my competitors?”
Answer: No. Just no. Not in your wildest, code-addled dreams. If I’m Google, which site will I prefer? The one that renders for me the same way it renders for users? Or the one that has to deliver two separate versions of every page?
When you spill banner ads all over your site, ask yourself…
You get the idea. Nofollow is better than follow, but banner pollution is still pretty dang far from perfect.
Mitigating SEO issues with search engine-specific tools is "fine." But it’s far, far from perfect. If search engines are forced to choose, they’ll favor the site that just works.
Not just SEO
By the way, distance from perfect absolutely applies to other channels.
I’m focusing on SEO, but think of other Internet marketing disciplines. I hear stuff like “How fast should my site be?” (Faster than it is right now.) Or “I’ve heard you shouldn’t have any content below the fold.” (Maybe in 2001.) Or “I need background video on my home page!” (Why? Do you have a reason?) Or, my favorite: “What’s a good bounce rate?” (Zero is pretty awesome.)
And Internet marketing venues are working to measure distance from perfect. Pay-per-click marketing has the quality score: A codified financial reward applied for seeking distance from perfect in as many elements as possible of your advertising program.
Social media venues are aggressively building their own forms of graphing, scoring and ranking systems designed to separate the good from the bad.
Really, all marketing includes some measure of distance from perfect. But no channel is more influenced by it than SEO. Instead of arguing one rule at a time, ask yourself and your boss or client: Will this move us closer to perfect?
Hell, you might even please a customer or two.
One last note for all of the SEOs in the crowd. Before you start pointing out edge cases, consider this: We spend our days combing Google for embarrassing rankings issues. Every now and then, we find one, point, and start yelling “SEE! SEE!!!! THE GOOGLES MADE MISTAKES!!!!” Google’s got lots of issues. Screwing up the rankings isn’t one of them.
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think of it as your exclusive digest of stuff you don't have time to hunt down but want to read!
This post was originally in YouMoz, and was promoted to the main blog because it provides great value and interest to our community. The author's views are entirely his or her own and may not reflect the views of Moz, Inc.
The spam in Google Analytics (GA) is becoming a serious issue. Due to a deluge of referral spam from social buttons, adult sites, and many, many other sources, people are starting to become overwhelmed by all the filters they are setting up to manage the useless data they are receiving.
The good news is, there is no need to panic. In this post, I'm going to focus on the most common mistakes people make when fighting spam in GA, and explain an efficient way to prevent it.
But first, let's make sure we understand how spam works. A couple of months ago, Jared Gardner wrote an excellent article explaining what referral spam is, including its intended purpose. He also pointed out some great examples of referral spam.
Types of spam
The spam in Google Analytics can be categorized by two types: ghosts and crawlers.
Ghosts
The vast majority of spam is this type. They are called ghosts because they never access your site. It is important to keep this in mind, as it's key to creating a more efficient solution for managing spam.
As unusual as it sounds, this type of spam doesn't have any interaction with your site at all. You may wonder how that is possible since one of the main purposes of GA is to track visits to our sites.
They do it by using the Measurement Protocol, which allows people to send data directly to Google Analytics' servers. Using this method, and probably randomly generated tracking codes (UA-XXXXX-1) as well, the spammers leave a "visit" with fake data, without even knowing who they are hitting.
Crawlers
This type of spam, the opposite to ghost spam, does access your site. As the name implies, these spam bots crawl your pages, ignoring rules like those found in robots.txt that are supposed to stop them from reading your site. When they exit your site, they leave a record on your reports that appears similar to a legitimate visit.
Crawlers are harder to identify because they know their targets and use real data. But it is also true that new ones seldom appear. So if you detect a referral in your analytics that looks suspicious, researching it on Google or checking it against this list might help you answer the question of whether or not it is spammy.
Most common mistakes made when dealing with spam in GA
I've been following this issue closely for the last few months. According to the comments people have made on my articles and conversations I've found in discussion forums, there are primarily three mistakes people make when dealing with spam in Google Analytics.
Mistake #1. Blocking ghost spam from the .htaccess file
One of the biggest mistakes people make is trying to block Ghost Spam from the .htaccess file.
For those who are not familiar with this file, one of its main functions is to allow/block access to your site. Now we know that ghosts never reach your site, so adding them here won't have any effect and will only add useless lines to your .htaccess file.
Ghost spam usually shows up for a few days and then disappears. As a result, sometimes people think that they successfully blocked it from here when really it's just a coincidence of timing.
Then when the spammers later return, they get worried because the solution is not working anymore, and they think the spammer somehow bypassed the barriers they set up.
The truth is, the .htaccess file can only effectively block crawlers such as buttons-for-website.com and a few others since these access your site. Most of the spam can't be blocked using this method, so there is no other option than using filters to exclude them.
Mistake #2. Using the referral exclusion list to stop spam
Another error is trying to use the referral exclusion list to stop the spam. The name may confuse you, but this list is not intended to exclude referrals in the way we want to for the spam. It has other purposes.
For example, when a customer buys something, sometimes they get redirected to a third-party page for payment. After making a payment, they're redirected back to you website, and GA records that as a new referral. It is appropriate to use referral exclusion list to prevent this from happening.
If you try to use the referral exclusion list to manage spam, however, the referral part will be stripped since there is no preexisting record. As a result, a direct visit will be recorded, and you will have a bigger problem than the one you started with since. You will still have spam, and direct visits are harder to track.
Mistake #3. Worrying that bounce rate changes will affect rankings
When people see that the bounce rate changes drastically because of the spam, they start worrying about the impact that it will have on their rankings in the SERPs.
This is another mistake commonly made. With or without spam, Google doesn't take into consideration Google Analytics metrics as a ranking factor. Here is an explanation about this from Matt Cutts, the former head of Google's web spam team.
And if you think about it, Cutts' explanation makes sense; because although many people have GA, not everyone uses it.
Assuming your site has been hacked
Another common concern when people see strange landing pages coming from spam on their reports is that they have been hacked.
The page that the spam shows on the reports doesn't exist, and if you try to open it, you will get a 404 page. Your site hasn't been compromised.
But you have to make sure the page doesn't exist. Because there are cases (not spam) where some sites have a security breach and get injected with pages full of bad keywords to defame the website.
What should you worry about?
Now that we've discarded security issues and their effects on rankings, the only thing left to worry about is your data. The fake trail that the spam leaves behind pollutes your reports.
It might have greater or lesser impact depending on your site traffic, but everyone is susceptible to the spam.
Small and midsize sites are the most easily impacted - not only because a big part of their traffic can be spam, but also because usually these sites are self-managed and sometimes don't have the support of an analyst or a webmaster.
Big sites with a lot of traffic can also be impacted by spam, and although the impact can be insignificant, invalid traffic means inaccurate reports no matter the size of the website. As an analyst, you should be able to explain what's going on in even in the most granular reports.
You only need one filter to deal with ghost spam
Usually it is recommended to add the referral to an exclusion filter after it is spotted. Although this is useful for a quick action against the spam, it has three big disadvantages.
Making filters every week for every new spam detected is tedious and time-consuming, especially if you manage many sites. Plus, by the time you apply the filter, and it starts working, you already have some affected data.
Some of the spammers use direct visits along with the referrals.
These direct hits won't be stopped by the filter so even if you are excluding the referral you will sill be receiving invalid traffic, which explains why some people have seen an unusual spike in direct traffic.
Luckily, there is a good way to prevent all these problems. Most of the spam (ghost) works by hitting GA's random tracking-IDs, meaning the offender doesn't really know who is the target, and for that reason either the hostname is not set or it uses a fake one. (See report below)
You can see that they use some weird names or don't even bother to set one. Although there are some known names in the list, these can be easily added by the spammer.
On the other hand, valid traffic will always use a real hostname. In most of the cases, this will be the domain. But it also can also result from paid services, translation services, or any other place where you've inserted GA tracking code.
Based on this, we can make a filter that will include only hits that use real hostnames. This will automatically exclude all hits from ghost spam, whether it shows up as a referral, keyword, or pageview; or even as a direct visit.
To create this filter, you will need to find the report of hostnames. Here's how:
Go to the Reporting tab in GA
Click on Audience in the lefthand panel
Expand Technology and select Network
At the top of the report, click on Hostname
You will see a list of all hostnames, including the ones that the spam uses. Make a list of all the valid hostnames you find, as follows:
yourmaindomain.com
blog.yourmaindomain.com
es.yourmaindomain.com
payingservice.com
translatetool.com
anotheruseddomain.com
For small to medium sites, this list of hostnames will likely consist of the main domain and a couple of subdomains. After you are sure you got all of them, create a regular expression similar to this one:
You don't need to put all of your subdomains in the regular expression. The main domain will match all of them. If you don't have a view set up without filters, create one now.
Then create a Custom Filter.
Make sure you select INCLUDE, then select "Hostname" on the filter field, and copy your expression into the Filter Pattern box.
You might want to verify the filter before saving to check that everything is okay. Once you're ready, set it to save, and apply the filter to all the views you want (except the view without filters).
This single filter will get rid of future occurrences of ghost spam that use invalid hostnames, and it doesn't require much maintenance. But it's important that every time you add your tracking code to any service, you add it to the end of the filter.
Now you should only need to take care of the crawler spam. Since crawlers access your site, you can block them by adding these lines to the .htaccess file:
It is important to note that this file is very sensitive, and misplacing a single character it it can bring down your entire site. Therefore, make sure you create a backup copy of your .htaccess file prior to editing it.
If you don't feel comfortable messing around with your .htaccess file, you can alternatively make an expression with all the crawlers, then and add it to an exclude filter by Campaign Source.
Implement these combined solutions, and you will worry much less about spam contaminating your analytics data. This will have the added benefit of freeing up more time for you to spend actually analyze your valid data.
If you still need more information to help you understand and deal with the spam on your GA reports, you can read my main article on the subject here: http://ift.tt/1Q0B4Ct.
Additional information on how to stop spam can be found at these URLs:
In closing, I am eager to hear your ideas on this serious issue. Please share them in the comments below.
(Editor's Note: All images featured in this post were created by the author.)
Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think of it as your exclusive digest of stuff you don't have time to hunt down but want to read!
I hadn t done much as far as effort to rank for terms around Houston SEO, that is until I posted this article. That was on July 24th, and I wasn t in the top 300 for any terms with Houston and Continued The post Finally getting ranked for SEO Houston appeared first on Houts Graphics. http://bit.ly/1SScLFs
I hadn’t done much as far as effort to rank for terms around Houston SEO, that is until I posted this article. That was on July 24th, and I wasn’t in the top 300 for any terms with Houston and … Continued