The MASC Word Sense Corpus

Year of Publication2012
AuthorsPassonneau, R. J., Baker C. F., Fellbaum C., & Ide N.
The MASC project has produced a multi-genre corpus with multiple layers of linguistic annotation, together with a sentence corpuscontaining WordNet 3.1 sense tags for 1000 occurrences of each of 100 words produced by multiple annotators, accompanied by indepthinter-annotator agreement data. Here we give an overview of the contents of MASC and then focus on the word sense sentencecorpus, describing the characteristics that differentiate it from other word sense corpora and detailing the inter-annotator agreementstudies that have been performed on the annotations. Finally, we discuss the potential to grow the word sense sentence corpus throughcrowdsourcing and the plan to enhance the content and annotations of MASC through a community-based collaborative effort.


This work was partially supported by funding provided to ICSI through National Science Foundation grant CNS: 0708952 (“Computing RES Infrastructure”). Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors or originators and do not necessarily reflect the views of the National Science Foundation.

Proceedings of the 8th Conference on International Language Resources and Evaluation (LREC 2012), Istanbul, Turkey, pp. 3025-3030

