Storieta
English
Save & sign up

About this book

The book is a straightforward compilation of common‑word lists for five major languages, French, German, Italian, Japanese, and Spanish, presented in plain ASCII text. It opens with a historical note explaining that the Ward word lists were among the largest public‑domain collections when they entered the Project Gutenberg archive in 2007, and it describes the technical format: phonetic spellings using backslashes to hint at accents, CR‑LF line endings, and a single zipped distribution for MSDOS. The introductory material also provides practical instructions for extracting the files, the size of each list, and a brief legal disclaimer about the use and redistribution of the text. In short, the volume is a utilitarian resource rather than a narrative, offering over half a million entries for researchers, programmers, or anyone needing a raw lexical dataset.

The tone is matter‑of‑fact and technical, reflecting the early‑2000s era of digital text archiving when Unicode was not yet standard. Its style is terse, with a focus on file specifications, licensing terms, and usage guidelines rather than literary flourish. Readers who appreciate raw data, need multilingual corpora for language‑learning tools, computational linguistics, or historical study of word‑list compilation will find it most useful. Those seeking a narrative or polished linguistic commentary may look elsewhere.

The opening · free to read

Historical Note:

The Ward word lists were some of the largest public domain word lists in the world, at the time they were added to the Project Gutenberg collection in 2007. These word lists do not contain 8-bit accented characters or Unicode, as would be found in a more recent Project Gutenberg eBook. Instead, the lists include phonetic spelling, utilizing backslashes and other characters to indicate where accents would normally occur. There is no detailed guide on how these extra characters were used, and therefore it is likely infeasible to map from the word lists back to a correct representation of the word (i.e., to map from a word list entry with slashes or other characters, back to the actual non-English word with accents or other non-ASCII characters).

These lists may still be useful, but they are no longer the state-of-the-art in word lists. In the time since the lists were created, it has become much easier for anyone with interests to make their own lists of unique words from the Project Gutenberg collection or other sources.

Moby (tm) Language II for MSDOS operating systems is compressed and distributed as a single zip file. After decompression the language files included with this product is in ordinary ASCII format with CRLF (ASCII 13/10) delimiters.

3) On the PG Catalog page click on the selection "More Files". You will see a "files.zip" folder in the list. Move this zipped folder to your computer. On your computer open "files.zip", double click on its "files" subdirectory and copy the contents into the destination directory on your computer.

Word lists in five of the world's great languages:

FRENCH number of words 138257 size in bytes 1524757 GERMAN number of words 159809 size in bytes 2055986 ITALIAN number of words 60453 size in bytes 561981 JAPANESE number of words 115523 size in bytes 934783 SPANISH number of words 86059 size in bytes 850523

Total number of words 560101 size in bytes 5928030

Once decompressed, the vocabulary files may be viewed and used just as any TEXT-type file might.

The book keeps going

Keep reading, and see it illustrated

Reading is free forever. Sign up and watch scenes appear while you read.

Illustrated scene from FrankensteinIllustrated scene from The Great GatsbyIllustrated scene from Pride and Prejudice

Scenes Storieta drew for other classics.

New illustrated classics

A new classic, drawn, in your inbox.

Once or twice a month: the latest books to get full character casts, scene art, and free comic editions. No account needed.