Showing posts with label duplicate. Show all posts
Showing posts with label duplicate. Show all posts

Thursday, September 17, 2009

New Update in Google Webmaster Tools: Submit URL Parameters to Ignore

I think this is definitely a good initiative by Google, allowing the webmasters to let the Googlebot know which URL parameters to ignore while crawling or indexing the site web pages.
  1. Log in to your Google Webmaster tools account
  2. Click on the ‘Site Configuration’ link on the left and then on ‘Settings’
  3. At the bottom you will find ‘Parameter Handling’ to adjust the parameter settings.

Dynamic parameters (for example, session IDs, source, or language) in your URLs can result in many different URLs all pointing to essentially the same content. For example: http://www.example.com/product?pid=123 might point to the same content as http://www.example.com/product. You can specify whether you want Google to ignore up to 15 specific parameters in your URL.

Also, Google lists the parameters they have identified in the URLs of your site and suggests if the parameters are vague for them. You can confirm your choice to ignore or not to ignore. You can also add parameters that Googlebot was unable to identify and list them.

The reasons why I find this useful are:

  1. This will really help Google in identifying the duplicate content pages and help them in proper indexing with fewer duplicate URLs.
  2. This can result in more efficient crawling of Googlebot.
  3. Googlebot can crawl more pages and index them, as the crawler efficiency is increased.
  4. This can reduce the PageRank dilution as external website may be linking to various versions of your URLs and this will help Google understand that all these pages are same. Thus, providing the proper PageRank credit to the page.
  5. Simple way to indicate the parameters to ignore or consider, instead of using the canonical option.

Yahoo Site Explorer has a similar feature and Bing is yet to include this. If you want to opt this method then you will be ignoring Bing and other search engines which do have the parameter suggestion option.

Related posts: Duplicate Content Issues and Probable Solutions

Tuesday, September 01, 2009

Duplicate Content Issues and Probable Solutions

I am sure many of the webmasters have this duplicate content issue and the Big G still does not has a clear-cut solution of solving this issue. Mat Cutts in his blog/videos did address this issue but, I do not think it completely solves the issue of giving the ranking credit to the actual owner of the article. How does Google know who the actual owner is?

Knowingly or unknowingly over the web duplicate content is developed and the SERPs are hijacked. So, let us know what leads to or is considered as duplicate content by Google or any other search engine.
  1. Treating mywebsite.com and www.mywebsite.com as different website. The simple solution for this is to use 301 re-direct for mywebsite.com to www.mywebsite.com.

  2. You want to have specific regional websites catering to different geographical areas. You have business.com and you also want to have business.co.uk, business.ca and so on. And you use the same content on all these domains. Matt Cutts says, it is not a big concern and webmasters need not worry about it. Do you have the answer for this?

  3. Issues with subdomains having the same content as the main domain. If at all you want to use subdomains then, it is better to use 301 re-directs.

  4. Dealers or affiliates taking the content from the website of main seller or the manufacturer to market and promote the products. These are honest people who want to sell the products and are caught as dupes as they usually copy the product description from the main website and have it on their own website. Sometimes the SERPs are also hijacked in such cases. If you want to get traffic from the search engines then, use your own content. Google always and will love “original” content.

  5. If you have an affiliate program then the program usually generate affiliate links similar to www.yourwebsite.com?affid=007. Search engine will consider this as a separate page to www.yourwebsite.com. Your affiliates will use the affid parameter links, so that they get the credit of sales/clicks their website generates. The only way to avoid this is ask your affiliates to use rel=”nofollow”. But, this can be a dumb idea of losing the link credit. But, this affiliate link can hijack your main site rankings.

  6. If your site is getting few pages dynamically with duplicate content like
    http://www.example.com/products.php?trackid=123
    http://www.example.com/products.php?reportid=123
    http://www.example.com/products.php?sessionid=123
    http://www.example.com/products.php?favouriteid=123
    http://www.example.com/products.php?print=yes&trackid=123
    Specifying a link tag in the section of your page content (as seen below) will indicate to the search engine robot that the URL is present and that it should be represented as the preferred canonical URL designated as: link rel=”canonical” href=”http://www.example.com/products.php” using proper html open and close tags. Please note that the tag will be treated similarly to the 301 re-direct.

Please note that this is just tip of the ice berg and still there are more such innocent duplicate content issues which are difficult to identify and resolve. So, always be careful you do not make such mistakes.