Your goal is to examine the “experimental corpus” models in the Sandbox and build a
forensic case for the specific properties or limitations in their preparation. If
you have a corpus of your own in the sandbox, you can also include that in your
investigation. Feel free to chat as a group and share notes and ideas as you
explore.
Look at two or three of the models: what can you determine about the choices that
were made in corpus preparation? What are the specific clues you can find by
exploring the models? Look for evidence related to one or two of the points below:
- Decisions about regularizing the language in the corpus
- Decisions about excluding or including paratexts
- Decisions about advance preparation of the language in the corpus to
combine up (tokenize) certain terms
Make a list of the evidence you can find about data preparation, and develop some
notes about how these data preparation choices seem to be impacting the models.
Some hints and things to think about:
- Try looking at clusters and see if you can find any clusters that are
revealing
- Try doing the same query on several different corpora (using the WWO full
corpus as a baseline): how do things like cosine similarity differ for
comparable words? What might this reveal?
- Try doing some of the operations, such as additions, subtractions, and
analogies: do they behave as expected?
Word Vectors: Hands-on Practice and Group Work, Intensive Research-focused slide 07
of 10
© 2019 Syd Bauman, Julia Flanders, Sarah Connell, and the Women Writers Project This
TEI-encoded XML file is available under the terms of the Creative Commons Attribution-ShareAlike
3.0 (Unported) license.