The first Project Gutenberg eBook was created on July 4, 1971. Michael S. Hart had been granted access to a powerful mainframe computer at the University of Illinois at Urbana-Champaign, and realized that his greatest impact would be by digitizing and distributing free literature (for more history, see: The eBook is 40 (1971-2011), by Marie Lebert, https://www.gutenberg.org/ebooks/36985).
Michael took a printed copy of the United States Declaration of Independence (www.gutenberg.org/ebooks/1) to the computer laboratory, where he sat at the teletype terminal and typed this first eBook. He distributed it via email to the people he knew about via the Internet’s predecessor, ARPAnet, which was available at UIUC. At that moment, the first eBook had been freely distributed to the online community of the day.
Digitization and production techniques, at the time of this first eBook, were /ad hoc/ and informal. A single eBook producer would edit a single file, from a single source. The first eBook’s printed source was a single sheet of paper, without hyphenation, a book cover, images, or other characteristics of book-length sources. In 1971, capitalization was not an issue, as only upper case letters were available in the character set used by the system.
Figure 1: Top view of a Model 33 Teletype, salvaged from the computer laboratory where Michael Hart typed the first eBook. The paper roll was where output would be printed.
During the next twenty years, from approximately 1971-1991, techniques of digitization would be dramatically improved, and regularized. Ongoing developments since then have tracked the available technologies for eBook creation and use, as well as preferences and interests of the many volunteers who would produce those eBooks.
Throughout the history of Project Gutenberg, these techniques, while refined and clearly articulated, have remained flexible (see the Volunteers’ FAQ at https://www.gutenberg.org/help/volunteers_faq.html).
Emphasis on the Public Domain
Project Gutenberg’s founder, Michael Hart, was motivated by completely free and unencumbered redistribution of literary works. Access to literary works enables literacy, which in turn opens the door to education and, it is hoped, opportunity. Interest in literary works that could be freely redistributed led to an emphasis on books and other items that are in the public domain.
The public domain is, today, understood to be those items that are not copyrighted. Copyright in the United States, where Project Gutenberg operates, is defined as a temporary monopoly by authors (or their agents), in order to benefit from commercial potential and thereby fostering continued creation:
“To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries” (United States Constitution, https://www.gutenberg.org/ebooks/5).
ITEMS ARE IN THE PUBLIC DOMAIN FOR ONE OF THREE REASONS
1. They are ineligible for copyright. In the US, this includes works created by the US Government;
Because of its emphasis on literary works, Project Gutenberg has mostly focused on items for which the copyright term has expired. Until 1998, this included items published 75 years earlier. For example, items from 1920 entered the public domain when their copyrights expired in 1995. The US Copyright Term Extension Act of 1998 changed the term to 95 years for most literary works, so new items (from 1923 onward) will not enter the public domain before 2019.
Figure 2: Michael Hart’s sunroom workspace in his Urbana home
There are over one million published works from 1923 and earlier, and these are the main items that Project Gutenberg continues to digitize and distribute. In addition, there were approximately one million works published in the United States from 1923-1964 but not renewed. Those items entered the public domain when their first copyright term ended, 28 years after publication. The copyright procedures utilized are online at https://www.gutenberg.org/help/copyright.html.
COLLECTION DEVELOPMENT POLICY AND EARLY MARKUP
The eBook collection, and all other aspects of Project Gutenberg, relies on volunteers to grow. Therefore, selection of items is done mainly by volunteers. Project Gutenberg seeks to limit duplication in the collection, and instead prefers to add items not already in the collection. Improvements to existing items is ongoing, mainly when errata reports are submitted by readers.
It took over two decades to release the first 100 eBooks, with #100 being published in 1994. Most of those first eBooks were collected through personal interaction with Hart. He would guide or participate in the digitization process, often developing procedures to deal with new characteristics. Footnotes and endnotes, italics and underscores, bold text, and different fonts all presented challenges for representation as plain text. Primitive markup techniques were developed, such as using an underscore character to surround underscored text, like this.
It was not until the mid-1990s that hypertext markup language (HTML) was first used, and at the time it was decided that Project Gutenberg eBooks should be wholly self-contained. A zip file would include all of the needed images, and external links were discouraged.
Throughout the entire history of Project Gutenberg, volunteers have been encouraged to work on items they are interested in, and to make their own decisions about how to best represent the content.
The book keeps going
Keep reading, and see it illustrated
Reading is free forever. Sign up and watch scenes appear while you read.
Scenes Storieta drew for other classics.