HathiTrust Releases Massive Open Dataset

HathiTrust has announced the release of a significantly expanded open dataset, HathiTrust Research Center (HTRC) Extracted Features (EF) Dataset, Version 1.0. This dataset provides researchers with open access to data extracted from the full text of the HathiTrust Digital Library (HTDL) at an unprecedented scale. 
 
The Extracted Features Dataset opens the complete HathiTrust collection for investigations into historical and cultural trends, the rise and fall of topics within the corpus, and the evolution of words and writing structures in publications dating from the 16th to the late 20th century. The EF Dataset provides quantitative information about word and line counts, parts of speech, and other details within each page of every volume in the HTDL. In addition to these larger-scale investigations, the EF Dataset also allows researchers to closely analyze the contents of a given volume or subset of volumes. The data is extracted from 13.7 million volumes found in the HTDL, representing over 5 billion pages consisting of over 2 trillion tokens (words). A preliminary release of the EF Dataset, drawn from a much smaller subset comprising only HathiTrust’s public domain collection, has already enabled novel research from scholars in economics, history, linguistics, literary studies and sociology, among other fields.
 
Read the full announcement here.

HathiTrust is a partnership of major academic and research libraries collaborating in a digital library initiative to preserve and provide access to the published record in digital form. UT Libraries became a member of HathiTrust in 2014. Read about benefits of UT’s membership in HathiTrust.