Can we use Google Scholar to identify highly-cited documents?

Martin-Martin, Alberto; Orduna-Malea, Enrique; Harzing, Anne-Wil; Delgado Lopez-Cozar, Emilio

Publicación: JOURNAL OF INFORMETRICS

2017

VL / 11 - BP / 152 - EP / 163

abstract

The main objective of this paper is to empirically test whether the identification of highly cited documents through Google Scholar is feasible and reliable. To this end, we carried out a longitudinal analysis (1950-2013), running a generic query (filtered only by year of publication) to minimise the effects of academic search engine optimisation. This gave us a final sample of 64,000 documents (1000 per year). The strong correlation between a document's citations and its position in the search results (r = 0.67) led us to conclude that Google Scholar is able to identify highly-cited papers effectively. This, combined with Google Scholar's unique coverage (no restrictions on document type and source), makes the academic search engine an invaluable tool for bibliometric research relating to the identification of the most influential scientific documents. We find evidence, however, that Google Scholar ranks those documents whose language (or geographical web domain) matches with the user's interface language higher than could be expected based on citations. Nonetheless, this language effect and other factors related to the Google Scholar's operation, i.e. the proper identification of versions and the date of publication, only have an incidental impact. They do not compromise the ability of Google Scholar to identify the highly-cited papers. (C) 2016 Elsevier Ltd. All rights reserved.

7_th in Computer Science

120 InfluRatio

14th Social Sciences, General